DocsThe record Architecture

Architecture

Collectors, the door, the job ledger, admission and retention.

How Elixir MCP is designed, for people who are curious and for anyone who needs to reason about what the service will and won't do. This page is the description of record — there is no fuller internal one, deliberately: a second copy is how the previous one came to describe a model that had already been replaced.

The shape of the system

Claude, ChatGPT, … Web interface explorer speaks the same tools elixir-bot MCP / OAuth email session service token Elixir MCP PostgreSQL S3 archive work queue — priority-ordered by Elixir MCP (collectors never choose targets) Collector 1 Collector 2 Collector n IP-bound keys · ONE shared rate budget, fleet = resilience Clash Royale API

Everything cloud-side is serverless (Lambdas + SQS + RDS Postgres, NAT-free VPC). The only machines with Clash Royale API keys are collectors — operator-run workers that lease fetch jobs from the queue, fetch with their IP-allowlisted key, and post results back. They never choose their own targets, hold no user data, and earn ladder points per call. The fleet shares one global rate budget by design: more collectors mean resilience, never more API load.

Collectors, in depth

A collector is a single static binary, Go or Python (its own public repo, elixir-mcp-collector) that runs anywhere with a static IP — a Mac in a closet, a Synology NAS at a cabin. What makes the fleet interesting:

  • Card identities. Every collector is named after a Clash Royale card — the operator's pick, and one collector per card — and appears publicly only by that card name; machine labels and IPs stay private. The console's Status page shows each card's heartbeat and hourly fetch rate, and /api/public/status publishes the same fleet without a session.
  • Credits. Fetches earn points, and points convert to the operator's own daily tool-call quota at 10:1 (capped at 4× the tier base). Running a collector literally buys your agent more questions.
  • Zero trust, zero AWS. A collector holds exactly two secrets: its operator's own IP-bound Clash Royale key, and a bearer token we issue. It speaks three HTTPS routes (config / lease / submit) — no AWS credentials, no queues, nothing that touches the tenant. The server computes each job's CR path and stamps the collector's identity onto results itself, so impersonation is structurally impossible and collection changes never require a client update.
  • Check-ins, not polling. A collector asks the door for work and is told when to come back (next_check_in_s: at once while work remains, otherwise the seconds until that collector's own slot in a fifteen-second idle cycle). Every active collector owns an evenly spaced slot, so a fleet of five idles at one check-in every three seconds instead of all five arriving together after each scheduler tick. The door never holds a connection open. Every collector serves priority work first, so a live: true fetch is picked up by whichever machine checks in next; there is no separate live channel.
  • Self-update, server-authorized. The config endpoint names the one binary version and SHA-256 a collector may install — a compromised release page alone cannot push code to operators. An update failure never stops collection; a dev build never self-updates.
  • Misbehavior is bounded. At most two unsubmitted leases at a time; leases that expire unsubmitted redeliver their jobs and count against the collector, and a collapsing submit ratio quarantines it automatically. Every payload is provenance-stamped forever, so even a late-discovered bad actor's data can be purged and replayed away.
  • Capture audit. Every fresh battle-log poll with prior coverage is audited: if the payload's oldest battle was previously unseen, the rotating log may have rolled past something — recorded as a possible gap. The 24-hour gap count is public on Status, so "no gaps" is a measurement, not a promise.

The timeline

Recording is pull; noticing is push. Everything you track appears on your timeline (elixir_timeline) while its notify switch is on. The timeline is synthesized when you read it, from the record and the per-subject ledger: the items that happened since your read pointer, in order, each a sentence a person can read with its facts beside it, and one entry per subject summarizing the window. A person's entries are the players they track and the clans they added; an agent's is the clan it represents, with the members inside it. Facts with their windows, never judgments: what to do about a member quiet six days is deliberately your agent's call, not the service's, and nothing on the timeline announces the time, which is game_clock's job. The shape lives on Timeline.

The feedback loop

Feedback is a first-class product surface, not a mailbox. Agents file it mid-session with elixir_feedback (attributed to the connecting account); people file it on the site. Every item gets a maintainer response — elixir_my_feedback pages through the full ledger, a feedback_responded event lands in your feed, and shipped fixes link the change. The same loop feeds the public Updates: contract versions are machine-readable (elixir_changelog), so an agent can ask "what changed since 0.20?" and discover capabilities that landed mid-session. Several shipped tools trace directly to agent-filed feedback.

The outbound relay

The VPC has no NAT — cloud components cannot reach the internet at all, which is a security posture worth keeping. The one exception is a small non-VPC relay Lambda fed by a queue. It does three jobs with deliberately different guarantees: transactional email (magic-link codes via JMAP — retried hard, dead-lettered loudly), anonymous Tinylytics product events (best-effort, dropped on failure), and newsletter enrollment at sign-in (Buttondown; idempotent, and an unsubscribed address is never re-subscribed). An analytics outage can never page anyone or delay a login email.

Recording: record once, entitle many

  • Canonical battles. The same battle appears in every participant's battle log; we store it once under a content-derived ID (time + sorted participant tags + battle class). Later observations enrich the record (fill missing fields) but never overwrite it.
  • Payloads and receipts. Every API response is content-addressed and receipted with which collector fetched it and when. Projections (battles, snapshots, war tables) are derived from payloads and are rebuildable from source.
  • The S3 payload archive. New payload content is archived to S3 at admission (Hive-partitioned by endpoint, entity, and date; gzip JSON; lifecycle to Infrequent Access at 30 days) before the database commit — a committed row always has its S3 twin. S3 is the system of record for raw observations; Postgres keeps a payload's JSON only for a fetch a reader is waiting on (a live: true request, read within seconds of admission) and only for two hours; an hourly sweep nulls it and retires superseded rows only after verifying their archived copy — the database holds metadata and projections, S3 holds the bytes. Athena (via one Glue table with partition projection) and DuckDB both query the layout directly.
  • Snapshots and events. Daily profile snapshots feed trophy/donation timelines; diffs between polls emit events with honest time semantics (most things are "observed between polls", and the data says so).
  • War data. Clan-scoped, multi-tenant (every war table is keyed by the observing clan), points-vs-fame discipline enforced at write time, and a per-clan war clock resolves battles to seasons/weeks/days from their own timestamps.

Scheduling: how often a player is fetched, and why

The official API is current-state only: a player's battle log holds roughly the last 30 battles and nothing older. Recording therefore means fetching often enough that no battle rolls off the log before a collector has seen it, while the whole fleet stays inside one shared, conservative API budget. The scheduler decides who is fetched when, from four rules that only ever shorten each other:

  • Observed yield sets the base cadence. Each recorded player carries an exponentially-weighted average of battles-per-hour actually harvested. The battle log is fetched when about five new battles are expected, clamped between 15 minutes and 24 hours, so an active player is polled tightly and a dormant one falls to daily on its own. A newly added player starts at an hourly discovery cadence.
  • A loss-aware bound stops bursts from rolling off. Yield learns from what a poll harvested, which is exactly what an overflowed log hides: a poll that returns 30 unseen battles after six hours reads as five an hour when the player may have played sixty. So at every admission the recorder also measures, from the battle timestamps themselves, the fastest this player has recently filled the log (the busiest six-hour window of the last 14 days), and polls before half that time has passed. A grinder who fills the log in two hours is fetched every hour for as long as that pace stays in their recent history; everyone else is unaffected.
  • A reader cap keeps the players you ask about fresh. Asking about a player through any tool marks that player as read, and their battle log then stays within an hour for the next day. Without this, a friend who plays a few games a day sits on the daily floor and can be a day stale precisely when you look. Reading costs nothing beyond the call itself; the extra fetches are a small, bounded slice of the budget.
  • A fairness floor guarantees no battle log is forgotten. Whatever the signals say, every recorded player's battle log is fetched at least daily.
  • The roster is the activity sensor. A clan roster carries the game's own lastSeen for every member in one small fetch. When a tracked clan's roster (read every 15 to 60 minutes while members play) is fresher than a member's last poll and shows they have not been in the game since it, that battle-log or profile poll is skipped: it would only return what the record already holds. A sighting younger than two hours never gates, so a session in progress is always followed. An incidental clan's roster is read every 4 to 24 hours and never gates: in the gate's first day those rosters held back the polls of players who then played a whole 25-battle session before the roster noticed, and capture gaps went from 0.1% to 1.5%.

Profiles are polled less often than battle logs - every eight hours once the roster shows a player active, with no floor for the idle, because the record keeps one snapshot per day and an idle player owes it none; the one time-critical profile read, the pre-reset capture of the weekly donation counter, is forced separately. Clan rosters follow the clan's own day — every 15 minutes while members of a tracked clan are in the game, coasting to hourly and then four-hourly as the roster's lastSeen stamps go quiet, and a few times a day for a clan read only because a recorded player is in it — and war-race polling reads the period type the API itself reports: tight on war days, relaxed on training days, with the war race also raising the battle-log cadence of members it names as having just battled. Because every cadence is a pure function of a player's state, subjects added together would otherwise fall due together forever; a stable per-subject phase offset de-phases such cohorts within a few cycles.

What this means when you read: every response carries the age of the polls it was built from (freshness_seconds in meta), elixir_coverage compares the lifetime battle counter against what was recorded over each observation interval and says so when they disagree, and the public status page reports how many of the last day's polls found the log had already rolled - as does capture_audit_24h in /api/public/status. When you need the state of play right now rather than the recorded history, live_fetch spends your live allowance on a fresh read instead of waiting for the schedule.

Access: entitlements, not permissions

  • A claim on your tag (trust-based; accounts are owner-approved) gives your agent your full history — including battles recorded before you joined.
  • Universal game reads: recorded player data — battles, profiles, timelines — is readable by every approved account, the same posture as the game's own public API. Clan cover gates the clan-scoped tools (roster, war): open members of a recorded clan use them, and that access ends the moment membership ends.
  • The MCP door is OAuth 2.1 with rotating refresh tokens. cr:read is the baseline, while recordings, collections, account preferences, and feedback each require their own write capability. The consent page names every requested capability, refresh never expands it, and an insufficient tool call is refused before it spends rate or daily quota.
  • Three principals, three doors. A person connects at /mcp; an agent at /a/<id>/mcp. Platform integrations use the REST API at /api/v1; /i/<id>/mcp remains a legacy migration surface. Grants are audience-bound to exactly one of them, so a credential presented at the wrong door is refused rather than quietly answering about the wrong subject — which matters because the three publish different tool surfaces and mean different things by "me". See Users, agents and integrations.
  • Service tokens are for headless principals — a Discord bot, a scheduled job, an app back end — and carry only the capabilities they were issued with. A human driving an agent signs in and consents instead, and never handles a raw key. The site uses cookie sessions; every tool call is audited against the credential that made it, with visible per-account quotas.

Honesty machinery

Coverage tools report recording start, per-endpoint freshness, and completeness ratios; timelines disclose their epochs; empty windows and inverted ranges refuse rather than pretending. The rule throughout: the system must never present a partial record as a complete one.

Operations

CI-gated public repo; collectors self-update from released builds; alarms route to an operations queue drained daily; performance is censused continuously (every ingest message logs phase timings). The database is audited periodically and raw payloads live in an S3 archive, content-addressed and queryable with SQL over S3. Site visits are counted anonymously with Tinylytics — see Privacy for exactly what that means.

The web API assembles feature-specific route modules behind shared session resolution. The account application separates overview, agents, connections, usage and feedback pages. The migration Lambda dispatches operational commands to separate modules; ordered migrations still run only through its deploy path. The shared metadata contract is checked when an envelope is built and after a tool returns. Successful UI journeys exercise the real API on scratch databases.

CI and local development use the same npm run verify gate: formatting, lint, dead-code analysis and workspace tests.

The code, and the family

Elixir MCP is built in the open:

All of it orbits POAP KINGS, the clan.