GPT, Claude, Gemini, and more

LLMFly AIYour models, behind one API.

Call GPT, Claude, Gemini, and more from the OpenAI SDK—without rebuilding your integration for every provider.

Want to test one request first?Read the quickstart

Let your AI handle the integration

Copy the integration guide into Claude Code, Codex, or Cursor.

View SKILL.md

Built for everyday API work

Spend less time maintaining model integrations.

Use one client, keep environments separate, and see the price before traffic reaches a model.

01

Separate development from production

Give each app and environment its own key. If one leaks, revoke it without taking the others offline.

02

Keep the client you already use

Change the base URL and choose a catalog model ID. Test provider-specific features only where your app needs them.

03

Know the rate before you send traffic

Compare input, output, and cache pricing in the catalog, then confirm each charge in usage history.

Quickstart

Make a working API call in three steps.

Create a key, point your client to LLMFly AI, and paste in a model ID from the catalog.

  1. 1
    Create a key

    Use a different key for each app and environment.

  2. 2
    Update the base URL

    Point your server-side OpenAI client to the LLMFly AI endpoint.

  3. 3
    Paste in a model ID

    Copy it from the catalog and start with a short, non-streaming request.

Where it works

Connect editors, agents, scripts, and backend services.

Use the same API endpoint from Claude Code, Codex, Cursor, automation scripts, or your own server application.

Claude CodeCodexCursorOpenAI SDKAnthropic SDKBackend APIs

FAQ

Questions developers ask before connecting.

What does LLMFly AI do?

It lets your server call the models in the LLMFly AI catalog through one OpenAI-compatible API.

Can I keep using the OpenAI SDK?

Yes. For compatible routes, change the base URL, API key, and model ID. Test tools, streaming, and structured output if your app uses them.

Which models can I call?

The catalog shows what is available now and which model ID to send. Availability can change.

How am I charged?

Each model has its own input, output, and—when available—cache rates. Your usage history shows what each request consumed.

Does LLMFly AI store prompts or responses?

Prompts and responses are processed to provide the requested API service. Retention varies by data type and is limited to what is reasonably needed for service delivery, security, disputes, and legal obligations; see the Privacy Policy for details.