Cost Architecture
Population scales. Inference does not. Only active interactions cost money.
$0.0275 per scene
3,500 input + 500 output tokens at $0.005/$0.020 per 1k. This is the core simulator profile, not a live provider quote.
Population Breakdown
Only active people (those in circles, with relationships, recently interacted) ever enter model context. 485 dormant people = $0 inference.
1:1 Conversation
20 messages
16 calls @ $0.0275/scene
3-Person Group
15 messages
12 calls @ $0.0275/scene
30-Message Evening
30 messages
24 calls @ $0.0275/scene
Road Trip (24h)
40 messages
32 calls @ $0.0275/scene
Proactive (10 candidates)
0 messages
No inference needed
Combined Core-Verified Examples
Core simulator profile — 500 people (15 active, 485 dormant)
Scenario totals use the core-verified token and rate profile. Population changes the dormant count, not these example workloads.
The Cost Ladder
CACHE
Identical scene context → instant return, $0
STATE / DB
Person, roles, relationships, circles — all zero inference
DETERMINISTIC
Relevance, selection, compilation, validation — pure code
PAID ROUTES
Remote provider classes are scaffolds. No real inference is connected.
CHATGPT
Disabled bridge contract. No ChatGPT inference or credits are connected.
GROKBOT
Disabled optional route contract. No GrokBot inference or credits are connected.
Why Population Does Not Drive Cost
Selection, Not Broadcast
ParticipantSelector picks ≤3 relevant speakers from circles. At population 5,000, 4,985 people never enter context.
Bounded Context
≤12k tokens, ≤30 sections per scene. Circuit breaker hard-trips on overflow. Context = constant.
One Call Per Scene
maxCallsPerEvent = 1. A 3-person group chat = 1 structured generation, not 3 separate calls.