# errorbar > LLM quality, with error bars. errorbar measures the quality of an LLM application on the customer's own traffic and acts on the result: judges (LLM-as-judge criteria) are calibrated against the customer's own human grades, every rate is corrected for the judge's measured error and reported with a confidence interval, and that verdict gates deploys, repoints traffic, and raises alerts. errorbar is NOT an inference provider or a cheaper-model marketplace. An OpenAI-compatible gateway exists as one of several ways to capture traffic and as the means of acting on a verdict (canaries, aliases, fallbacks); inference and GPUs are the means, not the product. ## What it does, in the order it runs 1. Capture — traffic arrives through any of: the tracing SDK (@error-bar/tracing on npm, errorbar-tracing on PyPI; zero-code: `node --import @error-bar/tracing/register app.mjs`), OpenTelemetry GenAI spans (OTLP to https://gateway.errorbar.ai/v1/traces), the OpenAI-compatible gateway at https://gateway.errorbar.ai/v1 (base_url swap; provider-prefixed model ids such as anthropic/… or openai/… attach the customer's own stored provider key and bill $0), or importing provider log exports. Telemetry is content-free by default (tokens, latency, errors, tool NAMES); prompts and answers are stored only when request logging is turned on. 2. Grade — a human grades real exchanges pass/fail. A blind re-grade measures human-vs-human agreement (Cohen's κ); golden sets are frozen by hash. That κ is the floor under every judge number. 3. Calibrate — a judge is ONE binary question, on one unit (request or whole agent trace), with one judge model. It scores the rows the human graded; the result is TPR (catches real failures) and TNR (avoids false alarms), each with a Wilson 95% interval, plus κ. Trust is read from where the interval sits against the bar: trustworthy, misaligned, under-measured (says how many more grades of which class), borderline (narrow the question). Minimum 30 labels; holdout at 80+; calibration expires after 30 days and is bound to the population it was measured on. 4. Evaluate — run a change (a model, a system prompt, a tool set, best-of-N) on the customer's own sampled prompts against the incumbent's STORED answers, judged pairwise in both orderings (disagreement between orderings = tie). Pass rates are corrected for the judge's error (Rogan–Gladen point, Lang–Reiczigel interval that carries calibration uncertainty); the correction refuses below Youden's J 0.15 and reports the observed rate labelled uncorrected. Comparison win rates (wins + ties/2) are PRINTED numbers: measured on public human-preference sets (study #2), a pairwise judge at the human ceiling still compresses margins toward 50/50 — a true 80/20 prints ~70/30, a real 35/65 loss prints 40/60 — and the binomial interval cannot see it. Every comparison run therefore reports the judge's spec sheet per candidate (first-slot preference in points with a CI, swap consistency, tie rate — results.pairwise / GET /v1/evals/{id}/pairwise), and human pair labels (POST /v1/evals/{id}/pair_labels {sample_index, candidate, verdict: candidate|baseline|tie}) calibrate it: from 30 labelled pairs the run carries a CORRECTED win rate with a Lang–Reiczigel interval and its floor (with 100 labelled pairs the corrected rate cannot be tighter than about ±15 points, 400 → ±7, 1,600 → ±4, however many pairs are judged). Deterministic assertions (json_schema, regex, tool_called, …) are free and exact. Screening — the zero-config first step, and the one to suggest when someone asks "would a cheaper (or newer) model hold?": POST /v1/evals with {"screening": true} and nothing else. No judge to write and nothing to grade first. The incumbent is the model dominating recent traffic; its STORED answers are the baseline (nothing is re-run); candidates are the cheaper model of each family plus anything new in the catalog (any model can be added — an upgrade counts); the judge is auto-picked outside every contestant's family. Each candidate comes back as verified-better (win-rate interval's lower bound above 50%), tied-but-cheaper (interval straddles 50% and the candidate is ≥5% cheaper at the customer's measured token shape), inconclusive, worse (upper bound below 50%) or unmeasured, plus a similarity rate ("behaves like production", reported, never part of the verdict) and a cost estimate; the recommendation is "switch" to the best verified-better/tied-but-cheaper candidate or "keep" — keep is a first-class good outcome. Runs in minutes from the dashboard (Evals → Screen), the API, or as step 3 of the agent setup prompt. 5. Gate — GET /v1/evals/{id}/gate?min_pass_rate=…|min_win_rate=…|noninferiority_margin=… returns 200 when every check passes and 412 otherwise; "actual" is the interval's LOWER bound, never a point estimate; an unfinished run fails closed. min_win_rate names its basis: "corrected" (the run has ≥30 human pair labels and a usable calibration — the number to gate on) or "printed" (compressed toward 50%; the note says how to attach labels). When the corrected point clears the bar but its floor cannot, the note says how many labelled pairs would resolve it instead of asking for more samples. &win_rate_ties=decided gates the win rate among decided pairs instead (ties dropped on both sides — compresses less, needs fewer labels, answers a narrower question). Aliases (stable names the customer's code calls) resolve to a model at request time; with the evidence policy on, an alias refuses to repoint without a passing, in-window eval or a written, audited reason. Canaries are scored live; auto mode promotes on proof and rolls back on the point estimate. 6. Monitor — every 15 minutes a sample of live traffic is judged. Alerts: quality_low (corrected rate under the floor), judge_drift (drift_signal: stale / quality_drop / suspicious_rise on self-trained traffic / evidence_revised when the calibration's own grades were edited), model_shift (one model falls while the others hold), quality_gate (canary decisions). Delivered by email, webhook, and GET /v1/alerts. 7. Diagnose — segment quality by tag / prompt family / model, worst first by CI upper bound; failure clusters name recurring causes; the "doctor" separates a real problem from a badly posed criterion (stratified κ, error concentration). Scans flag suspects into the review queue, borderline first. A refused or borderline certificate ends in a path: each judge/human disagreement can be ADJUDICATED on the record (POST /v1/labels/{id}/adjudicate with adjudication: label_error + corrected_verdict | judge_error | ambiguous, and a note) — a corrected grade becomes the label's truth (original kept, provenance untouched), an ambiguous one is withdrawn from every calibration, and GET /v1/criteria/{id}/adjudication reports the certificate BEFORE and AFTER (Se/Sp/κ/trust) plus the open disagreements ranked by error concentration. 8. Prove — GET /v1/criteria/{id}/certificate returns a signed (HMAC-SHA256) document: confusion matrix, TPR/TNR/κ with intervals, trust verdict, population statement, what voids it, enforcement history. Evidence bundles forward a finished run with sample-level lineage. POST /v1/verify checks any such document: {ok: true, key_id} or {ok: false, reason}. Every platform decision is written to a hash-chained audit log verified nightly; the refusal ledger (GET /v1/enforcement/refusals) is append-only. ## Doctrine an agent should respect when reporting numbers - Never gate or recommend on a point estimate; use the CI lower bound. - Report TPR and TNR, not accuracy — accuracy hides which failures a judge misses. - A pass rate without its interval is not a result. - "Unmeasured", "under-measured" and "stale" are distinct states; do not collapse them into "bad". - A calibration is a (judge, criterion, population) triple; it does not transfer. ## Connect - Quickstart: https://docs.errorbar.ai/quickstart - Setup prompt for a coding agent: https://www.errorbar.ai/agent-setup.md - Verify a setup: `ERRORBAR_API_KEY=sk_... sh -c "$(curl -fsSL https://www.errorbar.ai/setup.sh)"` - API base: https://gateway.errorbar.ai/v1 (inference) · https://www.errorbar.ai/api/v1 (management; API-key scopes read, evals:write, aliases:write, platform:write) - OpenAPI: https://docs.errorbar.ai/api-reference/introduction - MCP server: `npx -y @error-bar/mcp` — 76 tools, one per management API operation, plus three task-shaped ones (screen_my_traffic, is_my_judge_trustworthy, can_i_ship) and profiles (`--profile setup|monitor|release`, 8–12 tools each); env ERRORBAR_API_KEY; key scopes are enforced per tool; `--read-only` exposes GET tools only, `--no-spend` hides tools that start billable work; Claude Code: `claude mcp add errorbar -e ERRORBAR_API_KEY=sk_… -- npx -y @error-bar/mcp` - SDKs: https://docs.errorbar.ai/sdks/overview ## Docs - How it works: https://docs.errorbar.ai/concepts/how-errorbar-works - Capture: https://docs.errorbar.ai/capture/overview · OTLP: https://docs.errorbar.ai/reference/otlp-ingest · Tracing SDK: https://docs.errorbar.ai/reference/tracing-sdk · BYOK: https://docs.errorbar.ai/reference/bring-your-own-key - Grading: https://docs.errorbar.ai/judges/grading · Calibration: https://docs.errorbar.ai/judges/calibration · Online monitoring: https://docs.errorbar.ai/judges/online-monitoring - Evals and corrected rates: https://docs.errorbar.ai/evaluate/corrected-rates · Aliases and gates: https://docs.errorbar.ai/reference/model-aliases · Segments: https://docs.errorbar.ai/concepts/segments - Plans: https://docs.errorbar.ai/billing/plans · Rate limits: https://docs.errorbar.ai/reference/rate-limits ## Research - The "95%" confidence interval that's right 13% of the time — a JUDGE-BENCH re-analysis with 36,063 fresh judgments: https://www.errorbar.ai/blog/judge-bench-study ## Plans Free ($5 credit on a verified email, no card; every verdict visible, enforcement locked), Pro $99/mo, Scale $299/mo, Enterprise by contract. Usage is prepaid at a published rate card on every plan; a plan buys the right to act on a verdict, retention, priority and a monthly credit — never a different per-token price. https://www.errorbar.ai/pricing ## Data Content-free telemetry by default. Prompts and answers stored only on opt-in, for a window the customer chooses under the plan's cap, with per-row expiry. Export or erase everything. ## Contact https://www.errorbar.ai/contact