Integrate Claude API into Your Mobile App
Typical scenario: you're building a mobile chat assistant in Swift or Kotlin and choosing between OpenAI and Anthropic. Anthropic's Claude API offers up to 200k token context (claude-3-5-sonnet), native vision support, and excellent Russian language quality. We integrate Claude into your app turnkey: from backend-proxy architecture to final streaming testing. Our expertise — over 30 AI integration projects — helps reduce time-to-market by 2-3x compared to in-house development.
Secure Key Management
The Anthropic API key (sk-ant-...) must never be stored on the client. Rule: key stays on backend only. The mobile client communicates with your proxy server, which adds the x-api-key header and forwards requests to api.anthropic.com. Proxy architecture: any backend — Laravel, FastAPI, Cloudflare Worker. A minimal Cloudflare Worker implementation takes ~30 lines and handles both regular and streaming requests. Workers cold start — 5–10 ms, latency unnoticeable.
On the mobile client: JWT authentication to the proxy. The proxy validates the token, applies rate limiting (e.g., 20 requests/minute per user), and logs input_tokens/output_tokens for cost analytics. We ensure the key never leaves the backend.
Why Messages API Differs from OpenAI
The Anthropic Messages API differs from OpenAI Chat Completions in several ways:
- System prompt: separate
systemfield, not an element inmessages. Best practice: keep system context insystem, not inmessages[0]with role: "system". - Roles: only
userandassistant(nosysteminmessages). - No
function_calling— usetoolswithinput_schemain JSON Schema format.
{ "model": "claude-haiku-4-5", "max_tokens": 1024, "system": "You are a mobile app assistant...", "messages": [ {"role": "user", "content": "Explain this document"}, {"role": "assistant", "content": "Sure, ..."}, {"role": "user", "content": "What does clause 3 mean?"} ] } Streaming on Mobile: How to Speed Up Responses
Claude API supports SSE streaming with stream: true. The format differs slightly from OpenAI: content_block_delta event carries delta.text — that's one token; message_stop signals end of stream. On iOS, parse via URLSessionDataDelegate; on Android, use OkHttp EventSource. Delta events arrive every 10–50 ms during active generation. Buffer before UI updates: update @Published var streamText not at each event, but via a Throttle publisher (iOS) or distinctUntilChanged + debounce (Android Flow).
Comparison of streaming Claude vs OpenAI:
| Parameter | Claude (SSE) | OpenAI (SSE) |
|---|---|---|
| Event format | content_block_delta / message_stop | choices[i].delta.content / finish_reason |
| First token latency | ~350 ms (average) | ~300 ms |
| Vision in streaming | Yes | Yes (but via gpt-4-vision) |
| Client buffering | Throttle / debounce | Similar |
How to Analyze Images via Claude on Mobile?
Claude 3+ natively supports images in messages. Format:
{ "role": "user", "content": [ { "type": "image", "source": { "type": "base64", "media_type": "image/jpeg", "data": "<base64>" } }, {"type": "text", "text": "What is in this photo?"} ] } On mobile: compress image before sending. JPEG quality 70, max size 1568×1568 (API limit). Resize + compress via UIGraphicsImageRenderer (iOS) or Bitmap.createScaledBitmap + compress (Android). Token savings of 5–10x vs sending RAW.
Conversation Management and RAG
Claude handles 200k tokens, but for a mobile chat this is overkill and expensive. In practice, a sliding window of the last 20 messages is enough. For specialized apps (legal assistant, medical reference) — RAG (Retrieval Augmented Generation): store documents in a vector DB on the backend, augment the system prompt with relevant fragments per request. This doesn't grow history size but provides access to a large knowledge base. Learn more about RAG in Anthropic's documentation.
Handling Anthropic API Errors
529 Overloaded — servers overloaded, apply exponential backoff. 400 with error.type = "invalid_request_error" — usually max_tokens exceeded or invalid content format. 401 — wrong key on proxy. Log all errors with request ID (x-request-id) — needed for Anthropic support.
Case: legal assistant for a B2B app. Used claude-3-5-sonnet, contract analysis. User photographs a contract page, the assistant highlights key terms and risks. Image resized to 1200px on long side, JPEG 80. Average request: 2400 input tokens (image ~1800 + text 600) + 800 output. Streaming — first words appear in 350 ms. Users don't notice latency with streaming vs a "blank screen for 4 seconds" without it.
Claude Model Comparison
| Model | Speed | Context | Cost per 1M tokens (input/output) |
|---|---|---|---|
| Claude Haiku | Fast | 200k | $0.25 / $1.25 |
| Claude Sonnet | Medium | 200k | $3.00 / $15.00 |
| Claude Opus | Slow | 200k | $15.00 / $75.00 |
Model selection depends on the scenario: for simple chat use Haiku, for complex analysis — Sonnet or Opus. We help choose the optimal model and set up fallback to reduce costs.
What's Included in the Work
- Backend-proxy architecture (Cloudflare Worker / Laravel / FastAPI)
- Messages API integration with streaming and vision
- Conversation management (sliding window, RAG if needed)
- Token logging and error handling
- Deployment and support documentation
- Testing with TestFlight / Firebase App Distribution
Timelines and Cost
Basic integration with streaming, conversation context, and backend proxy — 3–5 business days. With image support and RAG — 1–2 weeks. Cost is calculated individually. Contact us for a project assessment — we'll prepare a proposal within a day. You can also order an architecture consultation — its cost will be deducted from the main project.







