MorphogenDocs

Anthropic SDK

Messages API, streaming, tools and prompt caching with the Anthropic SDK.

Morphogen accepts the Anthropic Messages format. Set base_url without /v1 and your Morphogen key. The format works for any text model in the catalog, not only Claude.

Connect

import os
from anthropic import Anthropic

client = Anthropic(
    base_url="https://api.morphogen.ru",
    api_key=os.environ["MORPHOGEN_API_KEY"],
)

The SDK adds the /v1/messages path and the x-api-key header itself.

Request

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=300,
    system="Отвечай коротко.",
    messages=[{"role": "user", "content": "Чем отличается TCP от UDP?"}],
)
print(msg.content[0].text)
print(msg.usage.input_tokens, msg.usage.output_tokens)

In this format max_tokens is required. Request parameters: Messages.

Streaming

with client.messages.stream(
    model="claude-sonnet-5",
    max_tokens=300,
    messages=[{"role": "user", "content": "Напиши хокку про дождь"}],
) as stream:
    for text in stream.text_stream:
        print(text, end="", flush=True)

Tool calls

tools = [{
    "name": "get_weather",
    "description": "Погода в городе",
    "input_schema": {
        "type": "object",
        "properties": {"city": {"type": "string"}},
        "required": ["city"],
    },
}]
msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=300,
    tools=tools,
    messages=[{"role": "user", "content": "Какая погода в Казани?"}],
)
for block in msg.content:
    if block.type == "tool_use":
        print(block.name, block.input)

Return the function result as a tool_result block in the next user message.

Prompt caching

Mark a long repeated prefix with a cache_control block:

msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=300,
    system=[{
        "type": "text",
        "text": open("instructions.md").read(),
        "cache_control": {"type": "ephemeral"},
    }],
    messages=[{"role": "user", "content": "Вопрос по инструкции"}],
)

The field passes to the model as is, and cache writes and reads are counted in the spending. GPT and Gemini models have an implicit cache, so cache_control is not needed for them.

Images in the input

import base64

data = base64.standard_b64encode(open("schema.png", "rb").read()).decode()
msg = client.messages.create(
    model="claude-sonnet-5",
    max_tokens=300,
    messages=[{
        "role": "user",
        "content": [
            {"type": "image", "source": {"type": "base64", "media_type": "image/png", "data": data}},
            {"type": "text", "text": "Что на схеме?"},
        ],
    }],
)

Errors

Errors come in the Anthropic format. On a 429, the SDK reads retry-after and retries the request itself. They have no error.code field, except for key_revoked: read the status, error.type and error.message. List of errors: Errors.

On this page