Tensic guides

Guide 03

Inference projects: connect any OpenAI- or Anthropic-compatible app

Create an inference project, connect an app with the OpenAI or Anthropic SDK, add embedding, image and speech models, set up model groups with failover, test in the playground and publish a project price.

On this page
  1. Create an inference project
  2. Connect an app with the OpenAI or Anthropic SDK
  3. Add embedding, image, speech and evaluation models to a project
  4. Create a model group for failover and zero-code model changes
  5. Test models in the inference playground
  6. Set a project price and publish it in the discovery file

An inference project is a stateless, OpenAI-compatible endpoint that serves the models you choose. Use it when you already have an app, SDK or tool that speaks the OpenAI API (or the Anthropic API). Point the app at Tensic and you get per-project budgets, API keys scoped to the project, failover between models, and logs of every call. Your code does not change.

Create an inference project

  1. In the side menu, open Projects and start a new project. You can also select + New Project on Home. Choose from scratch, or pick a template and select Back to Templates to return.
  2. Enter a Project name and choose the Team that owns the project. A project belongs to one team.
  3. Under Type, select Inference. The type cannot be changed after the project is created.
  4. Choose a Default model. This model answers when a caller sends no model, or names the default.
  5. Optionally, add Additional models. Callers can pick any of them per request with the model field. Only LLMs that the project's team can use are offered.
  6. Select Create project.

The project page has these tabs: Overview, Configuration, Orchestration, Access, Guards, Budget & Limits, Pricing and Logging. The Playground, Logs, Evals and API buttons are at the top. Overview shows tokens, requests, spend, the API keys scoped to the project, and the budget.

Every inference project always has a default model. Many OpenAI integrations do not send a model name, so the default model answers those requests.

Connect an app with the OpenAI or Anthropic SDK

Open the project and select API at the top. The API page shows ready-to-copy examples for this project, with your instance's URL already filled in.

API page of an inference project showing the base URL and an OpenAI SDK example

OpenAI spec. Use the official OpenAI SDK, or any tool built for the OpenAI API:

from openai import OpenAI

client = OpenAI(
    base_url="https://<your-instance>/projects/<project-id>/v1",
    api_key="YOUR_API_KEY",        # an API key scoped to this project
)

response = client.chat.completions.create(
    model="GLM-5.3",               # any LLM the project serves; see GET /v1/models
    messages=[{"role": "user", "content": "What can you help me with?"}],
)
print(response.choices[0].message.content)
  • The Base URL pins every call to this project. Copy it from the top of the API page.
  • model must be one of the project's models. An unknown name is refused with a 404.
  • The project's system message is always added on the server side, so do not send one yourself.
  • The page has Python, Streaming and cURL examples.

Anthropic spec. For clients that only speak the Anthropic Messages API, point the Anthropic SDK at https://<your-instance>/projects/<project-id>/anthropic. Leave out /v1, because the SDK adds it. Send the same project API key as x-api-key. max_tokens is required. The model can be any LLM the project serves; it does not have to be a Claude model.

Other endpoints on the same base URL:

  • Models (GET /v1/models) lists every model the project serves, grouped by kind.
  • Completions (legacy) (POST /v1/completions) accepts the older prompt format. Streaming is not supported on this endpoint.
  • Moderations (POST /v1/moderations).

The API page also has a one-line command that connects OpenCode to the project.

Add embedding, image, speech and evaluation models to a project

An inference project can serve more than chat models. Tensic covers most of the OpenAI API: text generation, embeddings, image generation, speech-to-text and text-to-speech. With that, an app built for OpenAI can generate images or transcribe audio through Tensic. You do not have to run your own GPU servers or image-generation software.

  1. Open the project and go to the Configuration tab.
  2. In the Model section, choose one of two options:
    • Turn on Serve every model the team has. The project then serves whatever the team is granted, including models added later.
    • Leave it off and pick the models for each kind yourself.
  3. For each kind, set a default and optional extras:
    • Default model / Additional models: LLMs.
    • Default embedding model / Additional embedding models.
    • Default image generator / Additional image generators.
    • Default speech-to-text model / Additional speech-to-text models.
    • Default text-to-speech model / Additional text-to-speech models.
  4. Optionally, set the Evaluation LLM, which is the model that scores eval runs. The default is the project's own LLM.
  5. Select Save.
Configuration tab, Model section, with default and additional models for each model kind

If a list says, for example, "The team has no other embedding models to add", your team has no access to more models of that kind. Ask your platform admin.

Create a model group for failover and zero-code model changes

A model group is a name that callers can send as model. Behind the name is an ordered list of models. The request goes to the first member that is available. If that member fails, the next member answers. Model groups give you two benefits:

  • Failover. If the first model is down or rate-limited, the next one answers.
  • Zero code changes. Your code asks for the group name, for example my-chat, instead of a model name. When a newer model comes out, change the group members in Tensic, and every app picks up the change without a deployment.

To create a group:

  1. Open the project and go to Configuration > Model groups.
  2. Select + Add group.
  3. Enter a Group name. Callers send this name as model.
  4. Under Members, in order, pick the models in order of preference. A group can only contain models that the project already serves.
  5. Select Save.

The usage log records both the group and the model that answered.

Model groups and Failover settings on the Configuration tab

Failover settings. The same tab has a Failover section that controls when Tensic stops sending requests to a failing model, and for how long:

  • Skip unhealthy models: when on, Tensic skips a model that is failing. When off, every request goes to the model it names, and the caller gets the real provider error.
  • Failures before skipping: the number of consecutive failures before a model is skipped. Leave it blank to use the instance default.
  • Cooldown (seconds): how long a model stays skipped after its last failure. Leave it blank to use the instance default.

Only provider-health failures count: a 429, a 5xx or a connection error. A wrong API key does not count, because waiting does not fix it. Tensic counts failures separately in each worker.

Test models in the inference playground

Every project type has a playground. To open it, select Playground at the top of the project.

  • Use the Model selector to choose which of the project's models to call. The default is preselected.
  • Use the tabs to test each kind of model: Chat, Compare, Image, Transcribe, Speak and Embed.
  • Stream and Auto-scroll control how the answer is shown. The live trace shows the timing of each call.

The inference playground works like a model playground: each message is a single, stateless call. It has no conversation memory and no tools. If you ask "what did I just ask?", the model cannot answer. This is expected for an inference project. For memory, tools and sessions, use an Agent project.

To see exactly what an app sends through an inference project, including its system message, open Logs at the top of the project. This helps you understand how an existing tool uses the model.

Set a project price and publish it in the discovery file

Each inference project can advertise its own price per model to the apps that use it. Coding agents such as OpenCode read that price to show how much each request costs.

  1. Open the project and go to the Pricing tab.
  2. For each model, enter the Input, Output and optionally Cached input price, in EUR per 1M tokens. Under each model name, the platform price is shown as "Platform: €x / €y".
  3. Select Save.
Pricing tab with per-model input, output and cached input prices

After you save:

  • The price appears in the project's discovery file, https://<your-instance>/projects/<project-id>/.well-known/opencode.
  • The price is returned as usage.cost on every OpenAI-compatible completion.
  • If you leave the fields empty, the API bills the platform price instead.
  • Leave Cached input empty unless the model charges less for a cached prompt. If it is empty, the whole prompt is billed at the input price.

The project price only changes what consuming apps see. It does not change what your instance is charged. You cannot edit the platform price.

The discovery file is public. The .well-known file needs no authentication. Anyone who can reach the project's URL can read it. By default, a model that the project has not priced is not published with the platform price. A platform admin can change this under Settings > General with Publish platform prices in the OpenCode discovery file. When this is on, models without a project price show the platform price, and a model that has no price at all is published as 0.

Common questions

Do I have to change my app to use Tensic?

Usually not. Set the base URL to the project's URL and use a project API key. Most apps built for the OpenAI API work unchanged. Features specific to OpenAI's own platform may not be supported.

What happens if my app does not send a model name?

The project's default model answers.

How do I switch to a newer model without redeploying?

Have your app call a model group name instead of a model name. Then change the group's members on the Configuration tab.

Why does the playground forget my previous message?

Inference projects are stateless. Each call stands alone. Use an Agent project if you need conversation memory.

Can I use the Anthropic SDK?

Yes. Point it at …/projects/<project-id>/anthropic without /v1, and send the project key as x-api-key. Any model the project serves can answer, not only Claude models.

Why does a failover not happen when my API key is wrong?

Only provider-health errors (429, 5xx, connection errors) count as failures. An authentication error is returned to the caller straight away.

Who can see the price I set?

Anyone who can reach the project URL. The price is in the unauthenticated .well-known/opencode discovery file, and it is also returned as usage.cost on completions.

Can I limit how much a project spends?

Yes. Each project has its own budget, and API keys are scoped to the project. Use the Budget & Limits tab.