SporeLabs — Product Specification
Single source of truth for SporeLabs: what it is, who it serves, pricing, user flows, and the service behind them. This file wins on any product conflict — including vs older notes in
AGENTS.mdor chat history.Status: Active (2026-08-11). Customer-facing name is SporeLabs everywhere. Repo/code paths may still say
edgeflow. Superseded Edgeflow-era drafts live indocs/archive/— do not implement from them.Drop-in launch contract:
SPORELABS_DROPIN_NOW.md— OpenRouter upstream, Kimi auto incumbent, pass-through billing, surface fidelity, atomic settle, high prove-gate. Implement the drop-in path from that file; this file still wins on product scope.
0. Handoff brief (read this first)
Product in one sentence
Point your app at SporeLabs and your LLM bill goes down every week without you doing anything. Auto traffic starts on Kimi K3 (via OpenRouter) so the app keeps working like a frontier API. We watch the shape of the work your traffic is doing, move it to cheaper models that score the same, train small models for the jobs that deserve their own, cut the calls and tool invocations that were never earning their cost, and prove every change on your own traffic before it serves a request. Customers forgive cost; they do not forgive premature downgrade.
The one-sentence promise (customer-facing)
Paste one line. Fund the account. Your system gets measurably cheaper and faster from there — we do the work, you keep the savings.
The hook
- Paste one line. Hand it to the coding agent you already use, or change
base_urlyourself. OpenAI- or Anthropic-compatible, so nothing else changes. - Fund the account. Prepaid dollars; the wallet is debited at our real upstream spend (pass-through, no markup for now). That is the entire bill.
- We go to work. Auto starts on Kimi K3. We shadow cheaper equals, specialize, and cut waste — continuously — and only cut over when the high prove-gate clears.
- Your bill falls and stays fallen. Every change is proven against your own traffic and re-checked forever. When unsure, we keep the incumbent.
- Optional: hosted agents for jobs that are a loop rather than a completion.
The four moves we make (customer-facing taxonomy)
This is the vocabulary for the product, the portal, and marketing. Use these four words; do not invent parallel jargon.
| Move | What it means | Why they care |
|---|---|---|
| Reroute | Same job, cheaper model that scores the same on their traffic | Instant savings, zero quality cost |
| Specialize | Train a small model for one recurring job; it beats the big one at that job | Cheaper and better; the moat |
| Trim | Remove work that was never earning its cost — starting with tool schemas the model never chooses | Savings a model swap cannot reach |
| Prove | Nothing serves real traffic until it beats the incumbent, and we keep re-checking | Makes the other three safe to automate |
Say Trim plainly in customer language ("we stopped paying for work that wasn't doing anything"). §5.0.2 has what ships and what does not — a signal we cannot price is a signal we do not raise.
Simple experience (UX law)
- One primary CTA per state.
- Onboarding is one pasteable line handed to their coding agent — the primary path, not a power-user shortcut. If a customer has to read a setup guide, the onboarding is broken.
- No model shopping as homepage. No "configure a train job" as the happy path.
- Show the work we did, not the machinery. Money and outcomes first; mechanism only when asked.
- Never make the customer feel they now have a system to operate. If a surface reads as a chore they inherited, it is wrong.
- MCP parity is first-class — other agents will often be the main user; dashboard is not privileged.
Decided essentials (pointers)
- Priority: drop-in API + wallet → the loop → one-line MCP onboarding → hosted agents → support bot; customer embed demoted. Details in §5.2; what is killed/demoted in §3.3.
- Pricing shape (decided 2026-07-27): one line item — tokens. Prepaid dollar wallet; frontier traffic is pass-through of our real spend (no markup); owned adapters S/M/L $0.20/$0.60/$1.50 per M; the loop is $0 — included for every funded account; support bot free. No subscription and no plan split (killed 2026-07-27) — anything that gates a capability on a plan is a bug. "Improve" survives only as the internal name of the loop. Details in §7.
- Ownership split: we own training data/eval design, the router and finetune lifecycle, waste detection, and agent proposals; customers may import datasets, review results, pin routes, accept or decline. The default is we act, they audit — not "we suggest, they operate." Details in §5.0.
1. Vision
SporeLabs is the replacement for "send everything to a frontier API forever" — without making the customer risk quality on day one. Point traffic at us and fund the wallet; Auto starts on Kimi K3 so their app keeps working. From that moment we are working on the bill: reading the shape of every call (never its content), proving cheaper equals in shadow, training small models for jobs that recur, cutting work that never earned its cost, and cutting over only when the high prove-gate clears.
Guiding question: What are you overpaying a frontier model to do?
Contrarian bet: most real jobs can run on smaller specialized models once we see the traffic, evaluate, and fine-tune. We do not force small when eval fails — keep the incumbent until a cluster earns its cutover.
Second bet: a meaningful share of frontier spend is not "the wrong model," it is work that should not be happening — a tool called every turn that never changes the answer, a retry loop nobody noticed, context hauled into every request and never read. We can see that from call shape alone, and no model swap recovers it.
Where the money comes from: token / upstream usage on the prepaid wallet (pass-through for OpenRouter; catalog rates for owned adapters) and hosted agent compute. Nothing else. We deliberately do not charge for the intelligence; it is the reason to send us the traffic. Markup on pass-through is open (§15.1).
The uncomfortable honesty: the loop lowers the bill we bill against. We accept that — the bet is that an account whose costs fall every week never leaves, and that we grow into the volume. Do not resolve this tension by quietly weakening the loop or by cutting over early.
2. Positioning
| We are | We are not |
|---|---|
| Drop-in API replacement that starts on Kimi K3 and costs less every week you leave it running | Dumb proxy / OpenRouter clone with only catalog routing and no prove-gate |
| A service that does the optimization for you | A dashboard that shows you optimizations you should go do |
| Your finetunes, from your traffic | Generic auto-router over public models only (OpenRouter Auto) |
| Platform that abstracts training (we design data + evals) | DIY fine-tune console / MLOps suite |
| A system that finds and removes wasted work, not just expensive work | Cost dashboard / FinOps reporting tool |
| Hosted agents sold from real work patterns | Multi-agent swarm marketplace |
| MCP-first (agents as primary clients) | Dashboard-only product |
| One bill: tokens / real upstream spend. Intelligence included | Subscription that gates features behind a plan |
| Edge/browser as showcase + future runtime story | Embed-first product |
Vs OpenRouter (honest): We use OpenRouter as upstream plumbing so the drop-in is frontier-faithful on day one. They route across a public catalog and show usage. They do not specialize your stack from your traffic, own training/evals for you, detect wasted work, or sell hosted agents from those patterns. That gap is the product.
Vs "just use a cheaper model" (honest): anyone can downgrade a model and hope. The product is the proof — we only move work that passes the high gate on the customer's own traffic, and we keep watching after cutover. When unsure, we keep the incumbent.
3. Goals
3.1 North star
A team pastes one line, funds a wallet, and never thinks about model selection again. Their cost per unit of work falls month over month, every drop is attributable to something we did and proved, and they did no work to get it. MCP works as well as the dashboard.
3.2 Definition of done (next product bar)
| Gate | Done when |
|---|---|
| Connect in one line | A customer can paste a single line into their coding agent and be serving traffic through us, with no docs read and no dashboard visit |
| Drop-in API | Account + wallet → Auto API key → OpenAI and Anthropic surfaces work (stream, tools, JSON); Auto serves Kimi K3 via OpenRouter; wallet debits real upstream spend; empty → insufficient-funds (SDK-mapped); see SPORELABS_DROPIN_NOW.md |
| Loop always on | Every funded account is audited, clustered, and optimized. No capability is subscription-gated. |
| Reroute | Work moves to a cheaper equal only after the high prove-gate (≥20 shared cases; strict agreement); incumbent fallback; auto-revert on continuous-eval fail |
| Specialize | Recurring patterns earn a trained model; we design data + evals and pay for the run (rudimentary NOW; polish later) |
| Trim | Wasteful tool calls / redundant work are detected and surfaced as a concrete fix, with the savings quantified |
| Prove | Nothing serves real traffic before beating the incumbent; continuous eval keeps checking after cutover |
| Attribution | Every dollar of reduced spend traces to a named change we made, on a date |
| Agents | From a pattern or intent: draft → persist → run on our cloud; metered compute |
| MCP | Agent clients can do the same flows as the dashboard (keys, wallet, insights, router, agents, evals) |
| Showcase bot | Free in-browser support bot works guest / no charge (edge story) |
Customer site-embed of their model: not a launch gate. If connect-in-one-line or the loop happy path fails once in a clean session, we are not done.
3.3 Explicitly killed / demoted
Killed: the Improve subscription and the Usage-vs-Improve plan split (2026-07-27); any paywall, upsell, price, or cancel flow for intelligence features; flat Stripe "Go live" subscription as the only money path; guest keyless customer embed without a funded account; model-tab hero / three base recipes as primary demo; customer-arbitrary HTTP tools at first agent launch; marketing that frames the loop as a process the customer runs — it is work we do for them.
Demoted (keep code / tell later): customer browser embed of production models as a primary CTA; manual train-tier purchase as the happy path (engine remains; UX is loop-owned).
4. Who it's for
| Persona | Goal | What we remove |
|---|---|---|
| Frontier API spender | Cut the bill without quality collapse or a project | Forever paying frontier rates for repeatable work |
| Agent builder burning tokens | Stop a loop from quietly costing $4k/mo | Hand-auditing traces to find the expensive turn |
| Drop-in team | Swap the endpoint today, get value passively | Migration project / rewrite |
| The busy owner | Not to become an ML person | DIY taxonomy, fine-tune ops, eval harness, model shopping |
| Agent buyer | "Run this job for us" | Framework + infra glue |
| MCP / their agent | Same outcomes via tools | Dashboard babysitting |
| Edge-curious (later) | Browser / on-device specialized models | Building runtimes themselves |
The disqualifying reaction: "so now I have another system to manage." Any surface that produces it has failed, regardless of how good the underlying analysis is.
5. Product surfaces
5.0 Killer experience — connect → we work → the bill falls
Headline product. Market SporeLabs as the API that costs less every week. Every stage below is on for every funded account — there is no plan column because there is no plan.
| Stage | What happens | What they see |
|---|---|---|
| Connect | One pasted line; OpenAI-compatible | Traffic flows in minutes; nothing else in their code changed |
| Read | Structure calls: task signals, volumes, latency, cost, failures, tool calls. Content discarded. | Nothing yet — this stage is deliberately invisible |
| Cluster | Group into named work types | "Here is what your app actually does," ranked by spend |
| Reroute | Move a pattern to a cheaper model that scores the same | "We moved X. Same results. −$N/mo." |
| Specialize | Design data + evals, fine-tune, serve behind the custom router | "We trained a model for X. It beats the frontier model at X. −$N/mo." |
| Trim | Detect work that never earned its cost — dead tool calls, retries, unread context | "We stopped paying for X. −$N/mo." |
| Prove | Shadow → cutover on merit → continuous re-check | Pass rates and savings, always attributable |
| Agentize | Propose a hosted agent for patterns that are really loops | Agent offer → accept → metered runs |
| Edge (later) | Same specialized models via browser/device runtimes | Story + support-bot showcase today |
Rules:
- Reading exists to act, not to surveil, and not to report. A finding we never act on is not a feature.
- Every stage is included. No capability may be gated on payment beyond a funded wallet.
- Training is abstracted: no happy-path "pick tier / upload JSONL / babysit job." We pay for the run.
- Customers may import datasets; handle securely (encryption, retention, delete/export).
- Router defaults are ours (Kimi incumbent when they omit
model); they may name a model on the request. - Do not force small when eval fails; keep the incumbent until a cluster earns cutover.
- Continuous eval watches quality and tool-call patterns over time; fail after live → automatic revert.
- No Project entity and no key-level auto/pin. Wallet and the loop stay account-scoped. Jobs are inferred from traffic. Keys are auth. Request
modelis auto/omitted vs a named id.
5.0.1 Attribution (how we prove we did something)
The product is only believable if every drop in spend has a name attached. For each change we make, persist and expose:
| Field | Meaning |
|---|---|
| Move | reroute / specialize / trim |
| Pattern | Which work it applied to |
| When | Date it went live |
| Evidence | Pass rate vs incumbent on their traffic |
| Savings | Measured $/month delta, not a projection |
The portal and MCP both render this as a history of work we did for you, newest first. This is the single most important surface in the product: the difference between "a tool I have to trust" and "a service that has already paid for itself."
5.0.2 Waste detection (Trim)
Surface a trim as one concrete change with a dollar figure, never as a report of anomalies. If we cannot name the fix and price it from measured data, we do not raise it.
Shipping today (platform/server/waste.py):
- Dead tool — a tool whose schema rides on nearly every request and which the model has never once chosen. We record the serialized size of each tool definition, so the waste is (schema tokens × requests) at the model's own rate. Measured, and the fix is one line.
Designed and deliberately not shipped, with the blocker for each:
- Retry loops. Needs a client request or session id.
shape_hashexcludes the user's text by design, so a busy pipeline sending one code path a thousand times an hour is indistinguishable from a client retrying. Shipping it would bill healthy traffic as waste. Returns when call events carry an idempotency or session key. - Unread context. Prompt size is measurable; how much of it was unnecessary is not. No ground truth, no dollar figure.
- Discarded output. We see one completion, not the caller's control flow, so "nobody read this" is not observable from the API surface.
5.0.3 The engine (one decider, not three features)
Reroute, specialize, and trim are three outcomes of one engine, not three subsystems a customer chooses between. platform/server/intelligence.py owns the decision; everything else is a capability it calls. Do not add a second decider.
Each tick, per account: re-audit traffic into patterns → advance whatever is already in flight → refresh waste findings → pick at most one new move on the most expensive pattern that clears the bar.
| Guard | Value | Why |
|---|---|---|
| Act on a pattern | ≥20 requests and ≥5% of spend | Below this the measurement is noise |
| Worth doing | ≥$0.20/mo projected | Not worth a customer-visible change |
| Moves in flight | 1 | The customer must always be able to name what we changed |
| Reroute attempts | 3 cheaper equals, then specialize | Bounded search before we spend GPU |
| Live cutover bar | The high prove-gate — SPORELABS_DROPIN_NOW.md §6.1 |
Thin suites must not move real spend; ties keep the incumbent |
Reroute first, specialize when reroute is exhausted. A cheaper equal is free to try and instant to revert; training is neither. The engine only opens a training run once the catalog has no untried cheaper model that holds quality on that pattern.
Auto incumbent: Kimi K3 via OpenRouter (EDGEFLOW_AUTO_INCUMBENT_MODEL, default moonshotai/kimi-k3). That answer set is the quality floor when the request omits model / sends auto.
Named-model requests: honor the id for that call. Same shadow comparisons on the inferred job; cheaper equals are suggestions until they stop naming a model. Never silent swap. Every call still feeds eval — they do not opt out by pinning.
Nothing reaches live traffic on a projection. Challenger and incumbent answer the same cases from the account's own traffic (eval_engine.run_comparison), and savings on the change record are recomputed from observed token mix at both models' real rates (pass-through or catalog). Continuous eval fail after cutover → automatic revert to incumbent.
Evals are minutes of inference, so they cannot run inside a request. Accounts enter a durable queue (loop_queue) on their first metered call; a worker Lambda on an EventBridge tick drains it with a per-account time budget. An unfinished eval is resumed, never restarted, and never blocks a completion.
Drop-in request path: one shared dropin_pipeline (OpenAI and Anthropic envelopes only), atomic settle in finally, OpenRouter for frontier ids, Modal for owned adapters. Contract: SPORELABS_DROPIN_NOW.md.
5.1 Free in-browser support bot (keep — it is the thesis, demonstrated)
The cheapest possible proof of the central claim: most of what you pay a frontier model to do does not need a frontier model. A ~1.7B model running in the visitor's own tab, with no key, no server, and no bill, answering real questions about the product.
- Landing demo, not the production loop. Helps them use SporeLabs / understand connect + the loop.
- Intentionally small; mistakes OK — say so tactfully, no apology tour.
- Fully free (not billed to grants). Prefer WebLLM in the browser; cheap hosted fallback may remain free (platform COGS).
- Frame it as evidence, not as a toy. The line to hit is "this answered you, and it cost nobody anything" — then let the visitor draw the conclusion about their own workload.
5.2 Delivery surfaces (priority)
- API (primary) — drop-in replacement; wallet; the loop included.
- MCP one-line connect — the fastest path from "interested" to "sending traffic."
- Hosted agents — we sell/run jobs on our hardware.
- Customer embed / on-device — demoted; future runtime story; support bot is the living showcase.
5.2.1 Onboarding (the one-line rule)
The primary onboarding is a single line the customer hands to the coding agent they already use. Their agent edits the config, sets the key, and confirms traffic. The customer reads nothing.
- The line must be copyable in one click and complete on its own — no placeholders except the key, and prefer flows where the agent fetches the key itself over MCP.
- It must work pasted into Cursor, Claude Code, or any MCP-capable client.
- A human-does-it-manually path stays available but is secondary, never the first thing on the page.
- Time from paste to first successful completion is a tracked metric (§14). Minutes, not a session.
5.3 Primary nouns
Happy path: traffic, patterns, savings, what we changed. Router, finetune, cluster, shape hash, and eval are mechanism — supporting detail, never the lead. Customers should feel they bought a bill that goes down, not "a fine-tune job" and not "an observability tool."
5.4 Hosted agent loop
Small, focused loop; host owns control flow; structured tools; compact errors. ~10 first-party tools at first ship (exact set = impl); no arbitrary customer HTTP tools at first ship. Encode 12-factor reliability; don't market the checklist. Prefer creating agents from Improve patterns (sell the job we already saw).
5.5 Evaluate (continuous, included)
- SporeLabs designs evals from traffic + synthetics; customer can add cases / import data.
- Continuous: quality, regressions, tool-call distributions, savings vs frontier.
- UI/MCP: pass rates + savings (disclose assumptions) — not MLOps theater.
- Eval exists to authorize a change, not to produce a report. Its customer-facing output is "this was safe to do, so we did it."
- Eval design and the training runs are on us. Generation/runner inference still meters wallet where we host it; keep that burn bounded (§15.6) and never let it exceed the savings it is chasing.
6. User flows
- Guest → belief (free): land; see the claim demonstrated (what the bill does over time and what we did to it); chat with the free in-browser support bot as live proof small models are enough; take the one-line connect instruction. No model-tab hero, no pricing table as the hero.
- Connect (the hook): copy one line → paste to their coding agent (primary) or edit
base_url(secondary) → sign in, top up → completions work immediately. Nothing else is asked of them, ever. The loop starts on its own. - The loop (automatic): traffic accumulates (features recorded, content discarded) → patterns named by spend → platform picks a move and proves it in shadow → cuts over on merit (or waits, per approval setting) → the change lands in attribution history with measured savings → optional agent proposals for patterns that are really loops. Approval posture is a setting, not a plan (§15.5); default biases toward acting automatically on reversible moves.
- Hosted agent: from pattern proposal or explicit intent; persist + operate needs account + wallet; runs meter compute; first-party tools only at first ship.
- MCP / their agent: connect MCP (
@edgeflow/mcpname may lag branding) — this is the one-line onboarding, not an advanced feature. Their agent takes an account from "nothing" to "sending traffic" without the human opening the dashboard. No tool is plan-gated; the dashboard is a peer UI. - Returning user: Auth0 Email OTP + refresh; Put it to work holds wallet, keys, and setup paths; Results holds spend, what we changed and what it saved, patterns, router, and agents. Stop topping up → usage hard-blocks at $0; the loop idles. No refunds, and no cancel flow — there is nothing to cancel. Leaving means pointing the base URL somewhere else.
State → primary CTA
| State | Primary | Secondary |
|---|---|---|
| Landed | Copy the connect line | Talk to the support bot |
| Signed in, no key | Create API key + top up | Hand setup to their agent |
| Traffic flowing, nothing changed yet | Open Results; show what we are reading and set expectation for the first change | Put it to work: keys / wallet |
| First change live | Open Results: see what we did and what it saved | Adjust approval setting |
| Steady state | Results over time | Accept agent proposal |
| Empty wallet | Top up (estimate) | — |
Note the absence of an upsell state. There is nothing to sell them after they connect.
7. Billing
7.1 Principles
- One line item: tokens. No subscription, no plan, no seats, no intelligence fee.
- Wallet = dollars for metered usage (inference, agent compute, any CDN/runtime we still bill). Prefer "dollars" in UI; "credits" OK only as casual synonym for wallet balance — not a second currency.
- The loop is included for every funded account: audit, patterns, reroute, specialization + its training runs, router, continuous eval, waste detection. We absorb the training and eval-design cost.
- Usage is priced near cost. We are buying retention and volume, not margin per token (§15.1 is open).
- Prepaid wallet by default. Auto-recharge opt-in; user-set amount; minimum $10. Optional hard monthly usage spend limit (default off). Spend trust tiers (Anthropic-style) before broad public abuse surface — ladder = impl.
- No refunds.
- Killed: the Improve subscription, the Usage/Improve plan split, and any flat sub that gates access or intelligence. If a capability check asks "have they paid for a plan," delete it.
7.2 Free (not billed)
| What | Amount | Notes |
|---|---|---|
| Support bot | Free | In-browser WebLLM only. Never bills the user. No anonymous API. |
| The loop | Free | Included with any funded account. Permanent property of the product. |
| Training runs for specializations | Free | We pay. Never billed to the customer wallet. |
There is no signup wallet grant and no anonymous API. New accounts start at $0. Inference requires an API key and a funded wallet.
7.3 Rate card
| Surface | Meter / fee |
|---|---|
| API chat/completions (OpenRouter / frontier) | Pass-through of our real OpenRouter spend (no markup) — millicents |
| API owned adapters (S / M / L) | $/M tokens (wallet) — $0.20 / $0.60 / $1.50 per M |
| Hosted agent runs | Compute (wallet) |
| The loop (audit, patterns, reroute, finetunes, router, eval, trim) | $0 |
| Specialization training runs | $0 — platform absorbed |
| Eval runner / generators | Platform-funded when proving a change; keep bounded (§15.6) |
| Customer embed CDN / edge runtime | Demoted; if used, bill honestly (CDN/compute) |
| Savings display | vs incumbent / frontier real rates + token estimates (disclosed) |
Under-the-hood train jobs may still use fixed internal tiers for capacity planning; do not make tier shopping the customer product.
Millicents: one small-model call costs a fraction of a cent, so the wallet meters in millicents (1/1000¢) and mirrors the familiar *_cents fields for display; rounding a call up to a whole cent would have erased the very savings we report. Accounts written before millicents are promoted on read, so no migration job is needed.
Frontier / Auto: debit OpenRouter-reported cost when present; else fail-closed estimate from tokens × maintained OpenRouter price table for that model id. Never skip the debit. See SPORELABS_DROPIN_NOW.md §4.
7.4 Empty wallet
Hard block on metered usage + top-up CTA with estimate. Auto-recharge uses configured amount (≥ $10) when enabled.
Drop-in SDK status mapping (stock clients):
| Surface | HTTP | Notes |
|---|---|---|
OpenAI /v1/chat/completions |
429 | insufficient_quota / insufficient_funds |
Anthropic /v1/messages |
400 | invalid_request_error |
| Intelligence "upgrade" paywalls | Bug | Never |
Do not return 402 on drop-in SDK paths — OpenAI clients mishandle it. The only legitimate refusal for metered inference is an empty (or hard-capped) wallet.
7.5 Auth & keys
- Named API keys (multiple per account; revoke individually) debit the wallet. Keys are auth, not routing policy.
- Omit
model/ sendauto→ we pick (Kimi until a cheaper equal earns live for that inferred job). Name a model on the request → that model for that call. - Auto copy: "Starts on Kimi K3. We only move traffic when a cheaper model matches on your traffic."
- Named-model copy: "We'll use this model for this call. We'll still measure cheaper options on the same work and show you what we'd change — we won't switch this request."
- No Project entity. No per-key auto/pin. Jobs are inferred from traffic.
- Empty → insufficient-funds (mapped per §7.4).
- Embed/site keys: keep for demoted/legacy embed path; not the hero. Legacy shared inference token: sunset on explicit ops timeline.
7.6 Entitlement
There is one entitlement question: is the wallet funded? If yes, everything works. There is no second question.
improve_plan.pyremains as the entitlement module, but every capability (insights,intelligence,router_serve) is on for every account. The capability names stay — MCP and the portal report them — but they no longer vary by payment.- Subscription lifecycle (
improve_status, grace windows,improve_subscription_id, Stripe subscription checkout and its webhooks) is dead. Remove it deliberately rather than leaving a dormant paywall that can be reawakened by a config flag. - Any
402/ "upgrade to Improve" response for an intelligence surface is a bug. Empty wallet on metered drop-in uses 429 (OpenAI) or 400 (Anthropic) per §7.4 — not a plan gate. - Imported customer data remains customer-owned; delete/export must be supported and is never gated.
- Historical accounts carrying
improve_statusfields are simply ignored on read. No migration job, no grandfathering path — we never charged for it at scale.
8. Website (portal)
8.1 Pages
Thin: marketing + workspace; account.html Auth0 callback. No chart maze. MCP connect linked prominently — it is onboarding, not documentation.
8.2 The show-don't-tell mandate
The site must demonstrate, not describe. Prose explaining a benefit is a failure state; the same claim rendered as a moving number, a falling line, or a working demo is the requirement.
Binding rules:
- No feature lists. A stack of bullets or a row of value-prop cards is prohibited on marketing surfaces. If information is genuinely enumerable, it belongs in a table of facts (prices, endpoints), not a list of claims.
- Every major claim needs a visual that proves it. Savings → a graph that descends. Cheap small models → a bot answering in-browser. One-line setup → the line, copyable, with the diff it produces.
- Show completed work, never a process diagram. Stage/pipeline graphics tell the visitor they have a system to operate. Show outcomes with our name on them: "We moved X. −$N/mo."
- Past tense and first person plural. "We trained a model for your ticket summaries" beats "specialization is applied to eligible patterns."
- Money is the unit. Lead with dollars; percentages second; tokens and pass rates only as supporting evidence.
- Never say "simulation." A worked example may be labeled honestly as an example, but framing that reads as a toy or a sandbox destroys the claim that we do real work. Disclose assumptions in a footnote, not a badge.
- Mechanism is available, not prominent.
shape_hash, shadow routing, and pass-rate thresholds are trust-builders for the skeptical reader — reachable one level down, never in the hero. - Clean hierarchy on mobile + desktop; one primary CTA per state.
8.3 Hierarchy
┌─────────────────────────────────────────────────────────────┐
│ SporeLabs [Sign in] / Wallet │
├─────────────────────────────────────────────────────────────┤
│ Your LLM bill goes down every week. You do nothing. │
│ [ one pasteable line ......................... Copy ] │
├─────────────────────────────────────────────────────────────┤
│ COST OVER TIME — a line that falls, annotated with the │
│ work we did to make it fall (reroute / specialize / trim) │
├─────────────────────────────────────────────────────────────┤
│ SUPPORT BOT — running in this tab, free, no key │
│ "this is how little a model needs to be" │
├─────────────────────────────────────────────────────────────┤
│ One bill: tokens. Everything above is included. │
└─────────────────────────────────────────────────────────────┘
8.4 Per-page job
| Page | Its one job | Must not |
|---|---|---|
| Home | Make them believe the bill falls on its own, then hand them the line | Explain architecture; list features |
Put it to work (usage.html) |
Fund the wallet, manage keys, and choose app or coding-agent setup | Mix setup with reporting or make users hunt across top-level tabs |
Connect your application (api.html, detail page) |
Show the drop-in SDK examples using a key from Put it to work | Compete in primary navigation or duplicate key management |
Agent setup (agent.html, detail page) |
One-time code + one MCP prompt for the coding agent | Duplicate the application setup page |
Results (improve.html) |
Money kept, visual evidence, finetune progress, suggestions with agent prompts, then spend, findings, proof, and completed work | Read as an invoice or a control panel of chores |
| Agents | Sell a job we already saw them doing by hand | Be a framework |
POST /v1/improve/refresh (and edgeflow_improve_refresh) accept an optional bounded focus string to steer the next eligible move toward matching pattern labels, task kinds, or tool names. Focus is transient — not stored, echoed, or audited — and never bypasses prove-gate, one-move-in-flight, or spend thresholds.
8.5 Models in UX
No primary base-recipe tabs. Advanced override only. The platform chooses bases/recipes for finetunes; the customer is never asked to.
9. Behind-the-scenes service
9.1 Architecture rules (do not break)
- No
*.modal.runin the browser — portal/accounts same-origin/api/*. - Modal = GPUs only (inference/train workers).
- Lambda IAM for accounts/DynamoDB — no AWS access keys in Modal for accounts API.
- Production deploy is GitHub Actions (merge to
main, or./scripts/ci/deploy-production.sh). Do not tofu-apply from a laptop — local.envis test/sandbox.
9.2 Component map
| Component | Role |
|---|---|
CloudFront sporelabs.dev |
Landing + /api/* + SDK |
| Lambda accounts API | FastAPI/Mangum — me, wallet, keys, Improve, patterns, router, agents, webhooks |
| DynamoDB | Accounts, tokens, jobs; extend for ledger, Improve state, call events, patterns, router, agents, runs |
| Auth0 | Passwordless Email OTP (SES); SMS legacy |
| Stripe | Wallet top-ups / auto-recharge only (subscription products retired 2026-07-27) |
| Modal | Inference + train GPU |
| S3 | Weights, datasets (imported + generated), manifests; secure handling for customer imports |
9.3 Runtime placement
- Hosted agents v0: short loops on Lambda or worker behind
/api/*; design for later durable pause/resume. - Never expose Modal URLs to the browser. Browser WebLLM: support bot showcase; future customer edge runtime is separate.
- Every platform-owned QLoRA run executes on Modal GPU functions. Trained GGUF artifacts are promoted through the shared Modal Volume; custom-model inference loads them in an
@app.clscontainer with@modal.enterand is reached only by the server-side inference origin.
9.4 Data model (direction)
| Concept | Meaning |
|---|---|
USER# |
Account, wallet, trust tier |
TOKEN# / API keys |
Auth + debit wallet |
EMBED# |
Demoted/legacy site keys |
| Call / usage events | Shape-only audit feedstock; no prompt/completion text; TTL'd |
| Bounded eval sample | At most 40 successful prompt/reply pairs per pattern, owner-scoped, 30-day TTL, owner-deletable; materialized eval cases/runs inherit the same TTL and are used only to prove/train that account |
| Pattern / cluster | Named work type |
| Router config | Pattern → model (+ fallbacks) |
| Change record | A move we made: type, pattern, date, evidence, measured savings (§5.0.1) |
| Dataset | Imported or generated; customer authz; delete/export |
JOB# |
Internal train jobs (platform-triggered, platform-funded) |
AGENT# / RUN# |
Hosted agents |
| Ledger | Usage debits/credits |
| Eval suites / results | Continuous + one-shot |
The change record is required — without a durable, queryable history of what we did and what it saved, the product cannot make its central claim.
10. API & MCP (first-class)
Inference stays completion-shaped at the model layer (drop-in): OpenAI /v1/chat/completions and Anthropic /v1/messages over one shared pipeline. Auto serves Kimi K3 via OpenRouter until a cheaper equal clears the high prove-gate. Endpoints: SPORELABS_API_REFERENCE.md; full drop-in contract: SPORELABS_DROPIN_NOW.md.
MCP is a primary client, and it is also the onboarding path. Package: @edgeflow/mcp (name may lag). Dashboard and MCP stay at parity for shipped surfaces. The pasteable line omits model (auto) unless they ask to name one.
Every capability below is available to every funded account. Do not reintroduce a plan dimension to this table.
| Capability | Available |
|---|---|
| Wallet / keys / me | ✓ |
| Usage stats | ✓ |
| Pattern insights / audit summary | ✓ |
| Savings history / what we changed | ✓ (§5.0.1) |
| Router get/patch | ✓ |
| Propose / apply specialization | ✓ |
| Continuous eval results | ✓ |
| Improvement loop logs | ✓ |
| Waste findings and their fixes | ✓ |
| Dataset import | ✓ |
| Agents create/list/run | ✓ |
| Internal train job APIs | platform-owned |
Do not require dashboard clicks for agent-driven workflows. An agent should be able to onboard the account, read the savings history, and approve a change end to end.
11. Auth & session
- Auth0 Passwordless Email OTP (SES; SMS/Telnyx legacy);
account.htmlcallback;localStorage+ refresh tokens (offline_access). auth0_audiencemust never default to client ID. Empty = SPA ID token./v1/mefailures must notlogout().- SES SMTP creds live in the Auth0 email provider config (console), not repo
.env.
12. Scope
In: everything in §3.2/§5 — drop-in API + wallet, one-line onboarding, the included loop, attribution history, Trim, continuous eval, secure data import, hosted agents, full MCP parity, support bot, clean portal UX.
Out / demoted: see §3.3. Additionally: multi-agent swarms; customer HTTP tool webhooks (first agent ship); refunds; native mobile apps; apologetic "we're small" marketing.
13. Launch gates
13.1 Foundation (mostly present — keep green)
- [x] Support bot works guest / no charge (WebLLM path)
- [x] Sign in with Email OTP → account; wallet starts at $0 (top up to run)
- [x] Inference requires an API key (no guest
/v1/complete) - [x] Top-up prepaid; balance debits on API infer
- [x] Empty wallet hard-blocks with estimate + top-up CTA
- [x] Named API keys (multi-key)
- [x] Hosted agent primitives exist
- [x] MCP package exists (extend for Improve)
- [x] No
*.modal.runin browser for accounts - [x] Portal branded SporeLabs
- [ ] Deploy path clean (
deploy_all.shdry-run /--executeas needed)
13.2 Killer path (built — keep green)
- [x] Call audit pipeline (privacy-safe) + rich stats (
call_features.py,call_event_store.py; features only, TTL'd) - [x] Bounded production-trace eval sample (
eval_sample_store.py; 40 pairs per pattern, 30-day TTL, owner deletion) - [x] Pattern clustering + insight UI/MCP (
pattern_engine.py,improve.html,edgeflow_patterns) - [x] Platform-owned finetune from patterns (no DIY train UX;
specialization.pyowns the job, customer wallet untouched) - [x] Custom router serving specialized models + fallback (
router_store.py, live inopenai_api.py; shadow until eval passes) - [x] Continuous eval (quality + tool-call trends + savings;
eval_engine.py, millicent metering so savings are real) - [x] Dataset import (secure;
dataset_store.py— encrypted, capped share of the training mix, export/delete) - [x] Agent proposals from patterns (
agent_proposals.py→ hosted agent) - [x] Customer embed demoted in IA/copy (embed endpoints remain; no portal or MCP surface leads with them)
Retired so no parallel product survives: customer train-job creation (POST /v1/jobs, /v1/jobs/generate-examples, /v1/train-tiers), the model-shopping page, and the DIY train tools in MCP and the hosted-agent runtime.
13.3 Repositioning gates (2026-07-27)
- [x] Improve subscription removed from entitlement, API, portal, and MCP; every funded account has every capability
- [x] Stripe subscription checkout / webhook / cancel paths deleted (top-ups untouched). Live webhook events are
checkout.session.completedandpayment_intent.succeededonly. - [x] Tests asserting Usage-plan denial replaced with tests asserting universal access, plus a test that the subscribe/cancel routes 404
- [x] Change record persisted and exposed (§5.0.1) —
change_store.py,GET /v1/changes,edgeflow_changes, and the Savings page leads with it - [x] Trim shipped: at least one waste signal detected, quantified, and surfaced as a named fix (§5.0.2) — dead tool schemas, priced per month, resolvable as done or dismissed
- [ ] One-line agent onboarding works from a cold start in Cursor and Claude Code (§5.2.1) — the line ships on Home and Connect; not yet verified cold
- [x] Portal/marketing rebuilt against §8.2: no feature lists, no process diagrams, no "simulation" framing, savings-first
- [x] Support bot retained and reframed as proof (§5.1)
14. Success metrics
| Metric | Target |
|---|---|
| Time from paste to first successful completion | Minutes, not a session |
| Share of accounts whose $/unit-of-work fell month over month | The north-star number. High and rising. |
| Time from connect → first change we made for them | Days, not weeks. This is when the product becomes real to them. |
| Cumulative measured savings per account | Rising; must exceed what they'd have paid a competitor |
| Accounts with a specialized model live | Rising |
| Trim findings shipped per account | Rising; each with a dollar figure |
| Continuous eval coverage on changed clusters | High |
| Changes cut over automatically vs waiting on approval | Automatic share rising — approval friction is failure |
| Patterns → accepted hosted agents | Quality > volume |
| MCP-originated onboarding share | Track; rising |
| Support load from billing confusion | Low (one line item makes this easy) |
15. Open questions (only unresolved)
Resolved items are not listed. Items 1–3 of the old list are resolved by killing the subscription. Remaining:
- Long-run monetization. Usage is priced near cost and the loop actively shrinks the metered volume we bill. Options not yet chosen: margin band on tokens, share-of-savings, paid hosted agents carrying the business, or volume economics alone. Do not resolve this by re-adding a subscription without an explicit human decision.
- Savings shown against their actual prior spend (needs a 30-day baseline) vs frontier list rates (available immediately). Prefer real baseline once we have it.
- How aggressive Trim should be about acting without approval — removing a tool call is more visible to their users than swapping a model.
- Audit retention, PII redaction, export/delete SLAs, and how imported data mixes with traffic-derived train mixes.
- Router v0 mechanics: shadow-vs-cutover default for specializations, routing mechanism (rules / classifier / embeddings), and default approval posture per move type (reroute and trim likely differ).
- Continuous eval cadence and wallet burn caps for eval traffic — must never cost more than the savings it chases.
- Exact first-party tool set (~10) for hosted agents; trust-tier ladder numbers.
- Rename public packages Edgeflow → SporeLabs timeline.
- Whether the change history should show changes we considered and rejected (builds trust, risks noise).
16. Ops & deploy
- Deploy: GitHub Actions on merge to
main, or./scripts/ci/deploy-production.sh. Localdeploy_all.sh --executetriggers those workflows. Never apply the accounts API from.env(test/sandbox Stripe). - Logs: CloudWatch + Modal. Customer audit payloads: stricter retention/access. Backups: DynamoDB PITR; dataset bucket policies.
- Secrets: AWS SM/SSM; Auth0 provider config for email (SES). Operator model deploys remain gated (
deploy_model.py --execute). - Repo runbooks (accounts, Stripe wallet, Auth0, domain, operator adapters) live in
docs/in the repo — internal, not published on this site.
17. Ownership map
File-level ownership lives in the repo's AGENTS.md (Quick orientation). Product-level split: portal in apps/edgeflow/ (brand SporeLabs; lead with savings + one-line connect); accounts/wallet + the loop + inference in platform/server/; MCP in packages/mcp; catalog/recipes in platform/products.yaml (advanced / internal).
18. Reference docs
Published on this site:
| Doc | Role |
|---|---|
SPORELABS_DROPIN_NOW.md |
Drop-in launch contract — Kimi auto, OpenRouter pass-through, surfaces, settle, prove-gate |
SPORELABS_JOBS_API.md |
Loop API + MCP surface (platform owns train UX) |
SPORELABS_API_REFERENCE.md |
Quick start + endpoint reference |
Internal ops runbooks (accounts, Stripe wallet, Auth0, platform engine, domain, operator adapters) live in docs/ in the repo but are not part of this site. This file wins on any product conflict.
19. Superseded decisions (do not revive without human ask)
| Old idea | New truth |
|---|---|
| Improve ~$20/mo for intelligence; Usage plan is wallet-only | No subscription and no plans. One bill: tokens / real upstream spend. The loop is included for every funded account. (2026-07-27) |
| Intelligence must remain gated on a plan | Gating intelligence is now a bug. Every capability is universal. |
Default serve = small starter model |
Auto starts on Kimi K3; cheaper models earn live only after the high gate |
| Soft prove-gate (thin case suites / loose agreement) | ≥20 cases, ≥85% agreement, −1pp pass tolerance, ties keep incumbent, auto-revert |
| Empty wallet → HTTP 402 on drop-in | OpenAI 429 / Anthropic 400 (SDK parity); 402 not used on drop-in paths |
| Markup on day-one frontier proxy | Pass-through, no markup until an explicit human decision |
| Project objects / per-key auto vs pin | Neither. Keys are auth. Request model is auto/omitted vs a named id. Jobs inferred from traffic. Wallet + loop stay account-scoped |
| Train tiers as customer-facing SKU | Training abstracted and platform-funded |
| OpenRouter feature parity as goal | OpenRouter is upstream plumbing for fidelity + the Kimi quality floor; differentiation is that we do the optimization work with a high prove-gate |
| Savings are the side effect of specialization | Savings are the product. Reroute, specialize, and trim are three means to it. |
Last updated: 2026-08-19. Canonical SporeLabs product spec — paste one line, fund a wallet, Auto starts on Kimi K3, bill falls only after proof. Pass-through upstream spend (no markup); high prove-gate; OpenAI + Anthropic drop-in; the loop included for every funded account; MCP-first onboarding. Drop-in contract: SPORELABS_DROPIN_NOW.md. Savings page leads with suggestions and proof; humans and agents share GET /v1/savings / edgeflow_savings. Wallet top-up minimum is $10.