LumenGuide
ANTHROPIC SDK AGENT LOOPS PROMPT CACHING
Service · Claude Integration

Production Claude apps — not chat demos.

We build with the Anthropic SDK every day. Tool use, the Agent SDK, prompt caching, batch APIs, citations, files, memory — all of it composed into hardened systems that ship to your production stack.

Scope a Claude build → Why Claude? ↓
agent.py · Claude Sonnet 4.6 CACHED
from anthropic import Anthropic
from aether import tools, cache

client = Anthropic()

response = client.messages.create(
  model="claude-sonnet-4-6",
  system=cache(system_prompt, ttl="5m"),
  tools=tools.crm + tools.calendar,
  messages=[{
    "role": "user",
    "content": query,
  }],
  betas=["context-management-2025"],
)

# cache_read: 8,420 tok · cache_write: 412 tok
# latency: 1.1s · cost: $0.0042 / req
Why Claude

When Claude is the right choice — and when it isn't.

We're not vendor-locked. We ship Claude when the task rewards it: long context, structured reasoning, tool use, and creative writing.

⧫

Long-context reasoning

1M token context windows on Opus/Sonnet — ingest entire codebases, contracts, knowledge bases.

⊞

Tool use that actually works

Native parallel tool calls, structured outputs, and a SDK that handles retries and budgets out of the box.

⊜

Prompt caching

5-minute and 1-hour caching can drop costs 70-90% on agent loops. We build with caching from day one.

⌬

Model routing

Opus for hard reasoning, Sonnet for default, Haiku for high-throughput — routed per task, not per app.

⌗

Citations + files

Built-in citation API and file handling — auditable answers from real documents, not hallucinated facts.

▢

Agent SDK

The Claude Agent SDK gives us memory, context management, and compaction without rebuilding plumbing.

Patterns we ship

From bot to autonomous agent — at the right altitude.

Most teams over-engineer their first AI feature. We start at the simplest pattern that solves the problem and only add agency where it pays for itself.

  • RAG chat over your knowledge base, with citations
  • Tool-use bots that call your internal APIs
  • Agent loops that plan, execute, and self-correct
  • Batch pipelines for high-volume classification & extraction
  • Coding agents that ship PRs against your repo
Scope your build →
agent.run · 3 tools, 2 hops claude-opus-4-7
→ planning…
thinking for 2.4s · 1,840 tok
→ tool: search_invoices(q="overdue", limit=20)
↳ 12 results · 230ms
→ tool: draft_reminder_email(invoice_id=…)
↳ 12 drafts · 880ms
→ tool: queue_for_review(drafts=…)
✓ queued in Outlook · audit log #4421
Real numbers

What "production-grade" actually means.

0 Avg cache hit rate
$0 Cost per cached call
0 p50 first-token
0 SDK retry success

Have a Claude prototype that needs to ship?

We pick up half-built Claude apps and harden them for production — eval suite, caching, observability, error budgets.

Talk to a Claude engineer →