Why Claude for a Mobile AI Assistant?
Imagine a user uploading a 100-page contract PDF and wanting to ask questions about it. Standard assistants with 8K context fail — you have to split the document, losing integrity. Claude solves this: its 200K token context window allows loading the entire document and answering without an RAG pipeline. We have over 5 years of experience developing mobile solutions and guarantee quality AI integration into your app. We evaluate your project in 1–2 days — just reach out to us.
Anthropic Messages API: Structure and Peculiarities
The Anthropic API is structurally similar to OpenAI, but with important differences. The system prompt in Claude is a separate system parameter, not a message with role system in the messages array. This is critical: trying to pass the system prompt inside messages degrades instruction-following quality.
struct AnthropicRequest: Encodable { let model: String // "claude-3-5-sonnet-20241022" let maxTokens: Int // mandatory, no default let system: String // system prompt — separate let messages: [Message] let stream: Bool enum CodingKeys: String, CodingKey { case model, system, messages, stream case maxTokens = "max_tokens" } } max_tokens in the Anthropic API is a mandatory parameter with no default. If you forget to pass it, the API returns a 400 error. This differs from OpenAI, where max_tokens is optional.
Authentication: the x-api-key header (not Authorization: Bearer). API versioning via anthropic-version: 2023-06-01. Without this header — 400 Bad Request.
How to Implement Streaming from Claude on iOS?
Claude supports streaming via Server-Sent Events. The stream structure differs from OpenAI: events content_block_start, content_block_delta, content_block_stop, message_delta — each carries its own fields.
Here is a step-by-step implementation on iOS:
- Initialize
URLSessionand create a request with headers. - Use
bytes(AsyncSequence) to read the stream. - For each line, check the prefix
"data: ". - Decode JSON into a struct with a
typefield. - For
content_block_delta, extract text and update UI on the main thread.
for try await line in response.bytes.lines { guard line.hasPrefix("data: ") else { continue } let jsonString = String(line.dropFirst(6)) guard jsonString != "[DONE]" else { break } if let data = jsonString.data(using: .utf8), let event = try? JSONDecoder().decode(StreamEvent.self, from: data), event.type == "content_block_delta" { let delta = event.delta?.text ?? "" await MainActor.run { self.appendText(delta) } } } It is important to handle all event types, not just content_block_delta — message_delta contains stop_reason (e.g., max_tokens), which you should show to the user.
Advantages of a Large Context on Mobile
200K tokens — roughly 150,000 words or ~500 pages of text. For a mobile assistant, this means working with full documents without an RAG pipeline. The user attaches a contract PDF — you can pass it entirely in the context and ask questions.
The downside: large context = long time-to-first-token. With 50K tokens in the request, the first response token can take 3–5 seconds even on a good connection. On mobile, you need a progress indicator that appears immediately, before the first token, otherwise the user thinks the app is frozen.
Cost also grows linearly with context — for apps with user billing, it is important to consider when designing a token counter UI. Claude 3.5 Sonnet processes 200K token context 2x faster than GPT-4o with the same volume, making it ideal for mobile scenarios with long conversations. Token savings can reach 40% compared to competitors, and overall infrastructure costs are lower thanks to native long-context support.
What Does a 200K Token Context Give to a Mobile User?
Let's compare key parameters of Claude 3.5 Sonnet and GPT-4o on mobile:
| Parameter | Claude 3.5 Sonnet | GPT-4o |
|---|---|---|
| Context window | 200K tokens | 128K tokens |
| First token speed (50K context) | 3-5 sec | 5-8 sec |
| Image support | up to 20, up to 5 MB | up to 10, up to 20 MB |
| Cost per million tokens (input) | significantly lower | higher |
| API structure | system separate, max_tokens mandatory | system in messages, max_tokens optional |
Claude wins on context volume and speed with large datasets. The API is stricter, but this reduces errors when configured correctly.
Vision: Sending Images to Claude
Claude 3.5 Sonnet supports images via base64 in a content block:
let imageContent = ContentBlock( type: "image", source: ImageSource( type: "base64", mediaType: "image/jpeg", data: imageBase64 ) ) Limitation: maximum 20 images per request, each up to 5 MB. On mobile, compress the image before sending to a reasonable size — UIGraphicsImageRenderer or BitmapFactory.Options with inSampleSize.
More on working with documents
For large PDFs, we pre-extract text via OCR libraries (e.g., PDFKit on iOS) to reduce token count. Alternatively, you can send multipart/form-data through a proxy server that strips extra headers.Process and What's Included
Key parameters to clarify: whether document support (PDF, images) is needed, expected conversation volume, whether a server-side proxy is needed (yes — mandatory, API key is not stored in the app).
What's included in the work:
- Documentation for Claude API integration on your platform
- Configured proxy server with authentication
- Ready Swift/Kotlin client for streaming
- Instructions for App Store Review (handling ATT, In-App Purchase requirements)
- 2 weeks of technical support after launch
Implementation: Anthropic API client → streaming UI → history management with 200K limit → optional file handling.
Estimated Timelines
Basic text assistant — 1–2 weeks. With document, image support, and server-side proxy — 3–4 weeks. Exact timelines depend on your stack and specifics. Request a free consultation — we'll show a demo version and calculate the cost. Over 50 AI integration projects under our belt. Contact us for a free project evaluation.
Anthropic API documentation: https://docs.anthropic.com/en/api







