Methods & decisions
How the numbers are made
The 2026 upgrade adds operations analytics, an A/B-test designer for the 15-minute late-discount rule, an append-only audit trail and an optional AI shift summary. This page says where the data comes from, which methods produce each figure, how they were checked, what they assume and where they fall short. The upgrade leaves the ported 2021 rules and their tests unchanged; the new work sits around them.
Data provenance
The original MongoDB database no longer exists, so everything on the site comes from a deterministic seed (web/src/db/seed.ts, PRNG seed 4399): the eight-item menu recovered from a page the 2021 app rendered, 15 van names from the team's original list (names only), 10 synthetic customers on reserved example domains, and three weeks of orders generated with the app's own pricing and order-id code. Minutes from order to ready are drawn uniformly between 4 and 21, so about a third of orders are late by construction. Live activity on the public site adds to that, and since it moved to a shared database it persists.
Consequence: the analytics demonstrate the methods; they are not findings about real vans. Full provenance is in the model and data card.
Operations analytics
/admin/analytics (one-click demo admin) computes, from every order in the database:
- Orders per day by Melbourne calendar day over 21 days, with the mean of complete days and a percentile-bootstrap 95% interval (2,000 resamples, seed 4399). Today is partial and excluded from the mean.
- Minutes from order to ready for served orders: a histogram, the median and 90th percentile with bootstrap intervals, and the share ready within 15 minutes with a Wilson interval.
- Time to fulfil as a Kaplan–Meier curve (shown as the share ready, 1 − S(t)) with Greenwood standard errors and log-log 95% intervals. Orders still being prepared are right-censored at their current age; cancelled orders are left out and counted separately.
- Housekeeping close-outs. A demo login closes out demo orders still active after 90 minutes and gives them an invented ready time so the order screens stay coherent. Those orders are flagged: their invented times are left out of every figure, the Kaplan–Meier curve censors them at their age when closed out, and the page states how many there were (DR-005).
- Late-discount rate by van: late served orders over all served orders, with Wilson 95% intervals, as a dot-and-interval plot so that small vans visibly carry wide intervals.
Every figure states its n. Every chart has a data table underneath.
Experiment design
/admin/experiments plans an A/B test of the late-discount rule: does promising a discount after 10 minutes instead of 15 bring customers back?
- Unit of randomisation: the customer. Each customer sees one rule throughout, and the analysis is per customer, so repeat orders from one person never count as independent evidence.
- One primary metric, chosen in advance: repeat order within 14 days, or a 4 or 5 star rating on the first order (a customer who does not rate counts as "no"). Its baseline is estimated with the same denominator: every non-cancelled demo order, rated or not, not only the rated ones. Guardrail: the share of first orders discounted (the cost of the promise).
- Sample size for two proportions with the pooled-variance normal formula, n1 = ((z1−α/2σ0 + zpowerσ1) / δ)², where σ0 uses the pooled rate under H0 and σ1 the two arms' own rates; Cohen's h (arcsine) as a cross-check. For a mean (stars), the two-mean formula with d = δ/σ, solved exactly for the z-test and with Guenther's correction for the t-test. All match statsmodels.
- Analysis of a simulated run: Wilson intervals per arm, Newcombe's hybrid score interval for the difference, Cohen's h and the relative lift as effect sizes, the pooled z-test, a seeded Monte Carlo permutation test (5,000 relabellings) and the exact permutation test (enumerated through the hypergeometric distribution). Every interval is built at level 1 − α for the α chosen, so the interval and the test answer the same question. The export carries the seed and the fulfilment pool, everything needed to rerun it.
- Peeking: 10,000 simulated A/A tests show how stopping at the first p < 0.05 inflates false positives, and how Pocock's group-sequential boundary restores the planned α. Simulations run in a Web Worker, and the peeking one is capped at 5,000 customers per arm, so a tiny minimum detectable effect cannot freeze the page; a zero effect has no sample size and is refused.
Evaluation design
The statistics are checked in two independent ways.
- Against reference software. Every estimator is pinned to values computed outside the app: statsmodels and scipy via
scripts/verify_stats.py(uv) and R's survival package viascripts/verify_km.R. Agreement is to 12 decimal places for the intervals, tests and Kaplan–Meier steps. - Against known truth. The experiment simulation injects a known effect, so its analysis can be scored.
pnpm calibrate(web/scripts/calibrate-simulation.ts) runs 2,000 simulated experiments per row at 583 customers per arm, baseline 35%, α = 0.05, and writesdocs/calibration.json, which this table is read from. The test suite re-runs both rows from the seeds and resampling pool recorded in that file and checks they match exactly; a faster 300-run check guards the simulation itself.
Calibration against known truth. Shares carry Wilson 95% intervals for simulation error; the mean difference carries its Monte Carlo standard error.
| Injected effect | Mean estimate | 95% interval covers the truth | z-test rejects H0 | Exact test rejects H0 |
|---|---|---|---|---|
| +8 pointsseeds 1 to 2,000 | +7.97 points (Monte Carlo SE 0.06) | 95.7% (94.7% to 96.5%) | 79.8% (78.0% to 81.5%)planned power 80% | 77.8% (75.9% to 79.6%) |
| 0 (A/A)seeds 2,001 to 4,000 | +0.01 points (Monte Carlo SE 0.06) | 95.4% (94.4% to 96.2%) | 4.6% (3.8% to 5.6%)nominal 5% | 4.0% (3.2% to 5.0%) |
Peeking (10,000 A/A experiments, seed 30005, 10 looks): looking once rejects 4.7% (4.3% to 5.1%), stopping at the first p < 0.05 rejects 19.6% (18.8% to 20.4%), and Pocock's boundary (|z| > 2.555) brings it back to 5.0% (4.6% to 5.5%). Resampling pool: 149 served orders from the seed snapshot (SHA-256 6c3d6e6d3c82…).
The 2021 business rules keep their own parity tests, which run the original JavaScript next to the TypeScript ports.
Audit trail
Order status changes (by vendors, customers cancelling, or the automated demo housekeeping), van open/close and location changes, sign-ins to the vendor and admin portals, CSV exports and AI review decisions are written to an append-only audit_log table in the same database batch as the change itself. SQLite triggers reject any UPDATE or DELETE on it; the only way rows leave is a full demo reset that wipes every table. It is visible in the records area with search and CSV export.
Assumptions
- Orders are independent of each other within a van and a day.
- Censoring of open orders is non-informative: an order still being prepared is as likely to finish soon as any other order of the same age.
- In the simulation, the rule changes the promise, not the crew's speed, and customers do not influence each other.
- The repeat-order baseline (35%) and 40 new customers a day are assumptions, not estimates: 10 synthetic customers cannot estimate them.
Limitations
- The data is synthetic; patterns reflect the generator.
- Per-van samples are small (4 to 44 served orders in the snapshot), so intervals overlap heavily: the data cannot rank crews.
- AI audit records are reported by the vendor's browser. The server checks the prompt, the output format and the fact check, but cannot prove a record matches a real provider response (DR-006).
- Cancellations are treated as a separate outcome, not as competing risks in the curve.
- The nearest-van ranking keeps the 2021 degree-space distance for parity, which overstates east-west distances by about 27% at Melbourne's latitude.
- Production storage. The public deployment ran on a separate copy of the database per serverless instance until 6 October 2026, so writes there were temporary (DR-004). It now uses a shared Turso database in Tokyo while the functions run in Sydney, and the seeded history is no longer re-dated, so it ages in place (DR-007). Pending schema migrations now apply to it on startup, and the seed command refuses to wipe it (DR-008).
What I'd change
- Connect the shared database before adding any feature that writes.
- Record the discount amount, so the guardrail can be measured in dollars.
- Model cancellations as a competing risk (cumulative incidence) instead of dropping them from the curve.
- Pre-register the experiment's analysis in the repository before any real launch, and add an always-valid sequential test for monitoring.
AI use statement
Snacks in a Van has one optional AI feature: a shift summary on the vendor's van page. Everything else in the app, including all analytics and the experiment designer, is ordinary code and statistics. The app works fully without any AI key, and makes no AI calls unless a vendor adds their own key and presses the button.
This statement and the controls it describes are informed by the Australian Government's policy for the responsible use of AI in government (Digital Transformation Agency), the transparency principles of the EU AI Act and the NIST AI Risk Management Framework. It is a portfolio demo and makes no claim of formal compliance with any of them.
What the AI does
- Turns today's aggregate figures for one van (orders, cancellations, sales, median minutes to ready, the share ready within 15 minutes with its interval, late discounts, the number and average of ratings, the top three items and the busiest hour) into a short note for the crew: a headline, a few highlights, watch-outs and one suggestion for the next shift.
- Runs only when the vendor presses "Write a summary", using the vendor's own Anthropic or OpenAI key.
- Default model: Claude Haiku 4.5 (lowest cost); Claude Sonnet 5.5 is offered for a stronger answer; for OpenAI the vendor types the model id.
What it never does
- It never changes an order, a van, a price, the discount rule or any statistic. Its output is text for a person to read.
- It never sees personal data: no customer names, emails, order ids or written comments. The figures are computed on the server from the orders and only the aggregates are put in the prompt (
computeShiftMetrics, with a test that the prompt holds none of those fields). The server also refuses to log any record whose prompt is not exactly the app's instructions plus a well-formed set of those aggregates, so a modified browser cannot slip other data into a logged call. - It never runs without a key and an explicit click.
- It never sends the key to this app's server. The call goes directly from the vendor's browser to the provider.
Data sent to the provider
Only the prompt shown under "Exactly what is sent to the provider" on the van page: fixed instructions plus the figures as JSON. The key travels in the request header to the chosen provider (Anthropic's Messages API, with the anthropic-dangerous-direct-browser-access header that browser calls require, or OpenAI's Chat Completions API), which processes the request under its own terms and bills the key's owner.
Where the key lives
In the vendor's browser only: session storage by default (gone when the tab closes, and cleared when the vendor logs out, because the demo vendor login is shared), local storage only if "remember on this device" is switched on (it then survives logging out), and removed by "Forget key". The key only moves between the two when the vendor flips that switch. It is never written to the audit log, never logged by the server and never committed to the repository.
Human in the loop and transparency
- Every output is labelled AI-generated with the provider and model. Text the vendor edited is labelled "AI draft, edited by the vendor", with the original draft one click away. Failed calls, which produced no output, are listed without the label.
- An automatic check compares every number in the output with the numeric figures that were sent and lists any that do not match, next to the text. Dates and clock times are accepted only in date or time form and only when they are the shift's date or the busiest hour; digits inside the date or a label never excuse a count. The server recomputes this check itself when the record arrives; the browser's own result is not trusted.
- The vendor records a decision: accept, edit (the edit is stored next to the original) or reject. Each output can be decided once. An edit that looks like an API key is refused, never stored.
- Before the browser contacts the provider, it asks the server for a short-lived, signed reservation; the server checks the vendor's session and the rate limits at that point, so a refused call never costs the vendor anything on their own key. After the call, every call made through the app, successful or not, is posted to the
ai_audit_logtable against that reservation: id, time, feature, provider, model, the input, the output (and the raw reply when it failed validation, was cut off or was refused), latency, token usage when the provider reports it, the fact check, whether the figures sent still matched the server's own, and the human decision. If the record of a successful call cannot be logged, the output is withheld; if the record of a failed call cannot be logged, the vendor is told. The server refuses anything shaped like an API key and takes the vendor's identity from the signed session. Decisions are also written to the append-onlyaudit_log. - The log is viewable at
/admin/ai-log(and as theai_audit_logtable in/admin/records) and can be exported as JSON or CSV. Exports are themselves recorded inaudit_log. - Database triggers keep each call record immutable apart from the single decision.
Known limitations
- Records are reported by the vendor's browser. The provider call happens in the browser (that is what keeps the key off this server), so the server receives a record of the call, not the call. It checks what it can (the reservation, the prompt, the output's format, the fact check) but cannot prove that a record matches a real provider response: the model name, latency, token counts and the reply itself are as the browser reported them. Because the demo vendor login is public, anyone could post a well-formed but invented record. The log shows what the app reported, attested by the signed-in vendor session (DR-006).
- The number check is a guard, not a guarantee: a summary can quote correct numbers and still mislead through emphasis or implied causes. That is why a person decides.
- Rate limits are kept in memory per server instance, so on a serverless deployment they are approximate.
- Since 6 October 2026 the public deployment writes the AI audit log to a shared Turso database, so records persist between requests and instances (DR-007). That cuts both ways: an invented record posted through the public demo vendor login persists too, until a full demo reset.
- Model outputs vary between runs and model versions; the log records which model produced which text.
Privacy & retention
Demo data only
Every customer, van and order in Snacks in a Van is synthetic, generated by a seeded script. Customer accounts use reserved example domains (example.com, example.org, demo.test). Every sign-up and sign-in form tells visitors not to enter real personal data, because accounts and orders are visible to anyone in the public records area.
What is stored
| Table | Personal data | Notes |
|---|---|---|
customers | Login id (an email-shaped string), first and last name, avatar choice | Passwords are bcrypt hashes and are never shown or exported |
orders, order_items, blogs | Linked to a customer login id; ratings and comments are free text | Comments are never sent to an AI provider |
vans, admins | Van names and an admin username; bcrypt passwords | Locations are public landmarks |
audit_log | The acting account's id (van name, admin username, customer login for cancellations, or demo-housekeeping) | No order contents, no free text from customers |
ai_audit_log | The van that made the call, and the aggregate figures sent | Never an API key; never customer-level data (the server refuses any record whose prompt is not the app's aggregate-only prompt, and any edit shaped like a key) |
Retention
- Locally (
web/data/app.db), data stays until you runpnpm db:reset, which deletes the file and reseeds. - On the public demo, since 6 October 2026 the database is a shared Turso database, so what visitors write (accounts, orders, posts, ratings, audit and AI log rows) stays until a full demo reset; there is no automatic purge yet (DR-007). A deployment without the database variables falls back to a copy of the seed per serverless instance, discarded when the instance recycles (DR-004).
- With a shared database, demo orders left active for more than 90 minutes are closed out at the next demo login (recorded in
audit_logasdemo-housekeeping, and flagged on the order so the analytics leave the invented ready time out). Nothing else is deleted automatically. - Audit tables are append-only. The only way their rows are removed is a full demo reset, which wipes every table at once. For a real deployment I would set an explicit retention period for both audit tables (for example 12 months), archive before deleting, and record each purge in a separate log.
Who can see what
- The records area (
/admin/records) shows every table to the demo admin, with password hashes redacted, and records every CSV export inaudit_log. - Vendors see their own van's orders. Customers see their own orders.
- The AI feature sends only aggregates to the provider the vendor chose; see the AI use statement.
Decision records
Each record states the decision first, then the context, the options, why, what actually happened (weak numbers included) and what I would change. Past records are never edited; a later one supersedes them.
- DR-001Rebuild the data layer on libSQL and Drizzle instead of MongoDBRead
- DR-002Replace Google Maps and OpenCage with MapLibre, OpenFreeMap and PhotonRead
- DR-003Port the 15-minute late-discount rule with three deliberate fixesRead
- DR-004Use Turso in production, with a /tmp copy as a read-mostly fallbackRead
- DR-005Treat housekeeping close-outs as censored, not as observed ready timesRead
- DR-006Reserve AI calls before they happen and re-verify their records on the serverRead
- DR-007Run production on the shared Turso database; keep /tmp for deployments without oneRead
- DR-008Apply pending migrations to the remote database on startup; refuse to seed itRead
Sources in docs/ on GitHub.