How to read this page: verifiability tiers

Not all evidence is equally strong, and an evidence page should say which kind it is offering. Everything on this page sits in one of three tiers, each labeled for exactly what it is worth:

  1. Self-measured, reproducible — every published number below. We measured them ourselves, so they carry our conflict of interest; in exchange, we publish the full recipe (frozen corpus, pinned checksums, exact commands, seeded runs) so anyone can re-run the measurement without trusting us.
  2. Independent verification — invited — what a third party has confirmed. Today: nothing. No third party has audited these numbers yet, and this page will not imply otherwise. A standing invitation to researchers and journalists is below.
  3. Planned — external evaluations we have committed to pursuing but that have not happened, listed without promised dates.

Tier 1 — Self-measured, reproducible

Everything in this tier was measured by us, on our own frozen corpus. Read it as a rigorous self-report: honest in method and fully disclosed, but not independent. The reproduction pack at the end of this tier exists so the claim “reproducible” is checkable, not decorative.

The frozen evaluation

Every number on this page comes from one measurement: engine v1 — the engine revision currently serving every scan — scored against a frozen, held-out evaluation corpus on August 16, 2026 (corpus manifest v5, seed 20260816). Every engine revision requires a fresh sanctioned measurement, so these numbers always describe the engine you are actually using. The frozen split is opened only by the gated measurement command — never during development — and every execution of that command is logged in an append-only ledger (disclosed below).

The evaluation set holds 1522 documents: 520 clean human documents, 528 clean AI documents, 160 human documents by non-native (ESL) writers, 310 paraphrase-attacked AI documents, and 4 further AI documents that belong to no named slice — self-generated samples that fall below the clean-AI slice's minimum word count, so they are scored in the run but excluded from the per-slice rates. It mixes clean, non-native and adversarial slices by design — a detector that is only measured on easy text publishes flattering numbers.

Verdicts are read at the pinned operating point: “likely AI” requires a calibrated probability of at least 0.93, “likely human” requires it below 0.50, and everything between is reported as inconclusive. Every rate below carries a Wilson 95% confidence interval — the honest range the true rate could sit in given the sample size, not just the point estimate.

Headline rates at the operating point

MeasureRateCountWilson 95% CI
False-positive rate — clean human writing (non-abstained) 0.2% 1 of 491 0.0% – 1.1%
False-positive rate — all human writing incl. non-native (non-abstained) 0.5% 3 of 609 0.2% – 1.4%
False-positive rate — non-native (ESL) human writing (non-abstained) 1.7% 2 of 118 0.5% – 6.0%
Detection rate (TPR) — clean AI text, all model eras (non-abstained) 82.4% 364 of 442 78.5% – 85.6%
Detection rate (TPR) — clean AI text, all model eras, abstentions counted as misses 68.9% 364 of 528 64.9% – 72.7%
Detection rate (TPR) — 2025–26-generation AI text (non-abstained) 99.1% 231 of 233 96.9% – 99.8%
Detection rate (TPR) — 2025–26-generation AI text, abstentions counted as misses 91.7% 231 of 252 87.6% – 94.5%
Adversarial recall — paraphrase-attacked AI text (non-abstained) 33.5% 73 of 218 27.6% – 40.0%
Adversarial recall — paraphrase-attacked AI text, abstentions counted as misses 23.5% 73 of 310 19.2% – 28.6%

“Non-abstained” rates are computed over the documents where the detector gave a verdict instead of saying inconclusive; the “abstentions counted as misses” rows charge every abstention against us. Both views are published because either one alone can flatter.

The “all model eras” clean-AI rows cover the entire clean-AI slice, which by construction mixes older-generation benchmark material (the RAID slice — see Datasets & licenses) with current output. The “2025–26-generation” rows are the same measurement restricted to our self-generated 2025–26 samples, aggregated from the per-family table below; current-era AI is easier for us to catch than the mixed-era headline suggests, and older-era AI is what drags the all-era raw rate down.

Ranking quality before the abstention band is applied (abstention-free AUROC):

SliceAUROC
Clean evaluation split0.952
Clean split including non-native human writing0.938
Paraphrase-attacked AI vs clean human writing0.855

What the verdict labels mean

The same three numbers the landing page shows, measured across the full evaluation set (clean, non-native and adversarial slices together):

OutcomeRateCountWilson 95% CI
Said “likely AI”, text was actually human 0.7% 3 of 442 0.2% – 2.0%
Said “likely human”, text was actually AI 26.9% 223 of 829 24.0% – 30.0%
Answered “inconclusive” instead of guessing 16.5% 251 of 1522 14.7% – 18.4%

Per-family detection on 2025–26 generations

Detection rates on the self-generated 2025–26 AI samples, split by the model family that wrote them. Small samples — mind the intervals.

Model family Docs TPR (abstentions as misses) TPR (non-abstained) Abstention
Fable 20 95.0% (19 of 20; CI 76.4% – 99.1%) 100.0% (19 of 19; CI 83.2% – 100.0%) 5.0% (1 of 20)
GPT 62 95.2% (59 of 62; CI 86.7% – 98.3%) 98.3% (59 of 60; CI 91.1% – 99.7%) 3.2% (2 of 62)
Haiku 90 91.1% (82 of 90; CI 83.4% – 95.4%) 100.0% (82 of 82; CI 95.5% – 100.0%) 8.9% (8 of 90)
Sonnet 80 88.8% (71 of 80; CI 80.0% – 94.0%) 98.6% (71 of 72; CI 92.5% – 99.8%) 10.0% (8 of 80)

Abstention rates per split

How often the detector says “inconclusive” instead of guessing, per evaluation slice. Abstaining more on hard slices is intended behavior — the alternative is confident error.

SliceAbstention rateCountWilson 95% CI
Clean evaluation split (human + AI) 11.0% 115 of 1048 9.2% – 13.0%
Clean AI documents 16.3% 86 of 528 13.4% – 19.7%
Clean human documents 5.6% 29 of 520 3.9% – 7.9%
Non-native (ESL) human documents 26.3% 42 of 160 20.0% – 33.6%
Paraphrase-attacked AI documents 29.7% 92 of 310 24.9% – 35.0%

The run ledger: how many times the eval has been opened

The frozen split loses its meaning if it is quietly re-run until the numbers look good. So every execution of the measurement command appends a row to an append-only, committed ledger (service/eval/EVAL-RUNS.md), and we disclose the count here: 6 recorded runs to date. Re-runs happened because the detection engine itself was iterated (each engine change requires a fresh sanctioned measurement, plus a determinism duplicate that verifies the run reproduces); the ledger records the stated reason for every row.

The reproduction pack

“Reproducible” is a claim like any other, so here is the full recipe — everything a skeptic needs to re-run every number on this page without trusting us.

Exact commands. From a checkout of the repository:

bash scripts/download_model.sh        # fetch + verify the pinned model
cd service
./venv-run python -m eval.fetch_data  # fetch + verify every corpus artifact
./venv-run python -m eval.run --reason "independent reproduction"

eval.run scores the frozen evaluation split with the exact serving scorer that powers /scan, writes service/eval/metrics.json (the file this page renders from), and appends its own row to the run ledger — including yours, in your clone.

Corpus manifest and fetch script. The corpus recipe is the committed manifest (service/eval/manifest.json): it pins every dataset artifact by URL, byte size and SHA-256 checksum, records the license and split assignment of each source, and states the split-hygiene rules (any single source dataset feeds at most one of train/dev/eval). The fetch script (service/eval/fetch_data.py) downloads the artifacts into a local cache and verifies every checksum — a mismatch that survives a re-fetch is a hard failure, never silent use. The frozen split ID lists are committed at service/eval/splits/, and our own 2025–26 generated samples are committed in full at service/eval/selfgen/.

Pinned artifact checksums. The sanctioned run records the SHA-256 of every artifact it depended on (repository state 9659213+dirty); a reproduction must be running against exactly these:

ArtifactSHA-256
Calibration artifact (score → probability mapping) ed5438ee01a3ba2abe6ac1d2bde282a2ec4caa9a30cea6a3a1807ca1ba042347
Score combiner c53c7aab91b8e8b54de80c3d02c288357442ee9a2765a7916d8ab74efa3111a1
Frozen evaluation split ID list bd6cb9ce7e5fb40efbd90ea43e3ae3ae0a0aab4b62da9cb6ac1e7f07613bb468
Signal-fusion configuration 6d3fa059406bfc6d1ba8ef61652bfc0223b53b51cf0e970ce64fbb207916cf1f
Corpus manifest d0de7265e176ee22d73962659766793f9c238398fe15906806e4c0b95ed152f3
Learned-classifier model artifact 6aa37981eeb049d83437f95738eea29d597ccddf470334d10e2113824291967b
Classifier training-corpus derivation dc62dcffe93fa529d684c5982d0f11189d4812109454bba9d54d1548cb86b047

Determinism guarantee. The measurement is seeded (seed 20260816, fixed in the manifest) and the pipeline is deterministic end to end: the run ledger records a determinism duplicate alongside each sanctioned measurement, verifying that a repeat execution reproduces the committed numbers. If your re-run of the commands above, at the pinned artifact checksums, does not reproduce the tables on this page, that is a finding — we want to hear about it.

Datasets & licenses

The evaluation corpus is assembled from published corpora plus our own generated samples; raw dataset artifacts are never committed — they are fetched from pinned URLs and verified by checksum. In summary:

  • Human writing comes from corpora with pre-LLM provenance: the Ghostbuster human collections (student essays, r/WritingPrompts stories, Reuters newswire; CC BY 3.0), the classic BBC News corpus (2004–05 articles; research-benchmark use, article copyright with the BBC), and — for the non-native slice — the PELIC learner corpus (CC BY-NC-ND 4.0).
  • AI writing comes from the RAID benchmark slice (MIT license), covering older-generation models plus DIPPER-paraphrase attacks, and from our own committed 2025–26 samples generated across four current model families (Haiku, Sonnet, Fable, GPT), including deliberately humanlike and evasive styles.
  • Training and development draw on further published corpora (HC3, CC BY-SA 4.0; MAGE; ELLIPSE, CC BY-NC-SA 4.0; a 2021 Wikipedia snapshot, CC BY-SA 3.0) — kept strictly separate from the frozen evaluation split.

Tier 2 — Independent verification: invited

No third party has audited these numbers yet. Everything above is self-measured, and a skeptical reader is right to weigh it accordingly — a detector grading its own homework has a conflict of interest no amount of methodological rigor fully removes. Independent verification is the strongest form of evidence a page like this can carry, and this page does not have it. We would rather say that plainly than let careful presentation imply a level of proof that does not exist.

So this is a standing invitation. To any researcher, journalist, educator or auditor who wants to verify — or falsify — our published numbers, we will provide, on request and at no cost:

  • the full evaluation corpus manifest with every pinned URL, byte size and SHA-256 checksum, plus the fetch-and-verify tooling;
  • the frozen split ID lists (train/dev/eval), so you can confirm the evaluation set we scored is the one we say we scored — and that nothing in it touched training;
  • the committed engine artifacts with their recorded checksums, and our complete set of self-generated 2025–26 samples;
  • the run harness itself, and help getting it running — including on an evaluation design of your own that we have never seen.

Write to research@cobalynx.com. We commit to publishing the outcome of any completed independent evaluation on this page — favorable or not — and to linking it from this section, which stays honestly empty until then.

Tier 3 — Planned: public leaderboard submission

Do not take our word for the field's difficulty — or our place in it. The RAID leaderboard (Dugan et al., ACL 2024, arxiv.org/abs/2405.07940) evaluates public detectors across models, domains and adversarial attacks, and is run by researchers with no stake in any detector. Our adversarial split reuses RAID's attack material precisely so our published recall is comparable in spirit; we have not yet submitted Cobalynx to the public leaderboard. Submitting is on our roadmap as a commitment, and when the entry exists we will link it here and move it up to tier 2. We list it as planned — not as done, and not with a promised date.

Context worth knowing when reading anyone's accuracy claims: the EU Code of Practice on AI-content transparency (July 2026) concluded that forensic detection of unwatermarked AI text is “not yet considered reliable enough,” and independent audits routinely measure commercial detectors far below their marketing numbers. Our methodology page explains how we operate inside that reality instead of pretending it away.

How we compare on transparency

We name no competitor here. Instead, here is the checklist we believe any detector's evidence page should be held to — including ours. Take it to any product you are evaluating and ask the same questions. This table compares publishing practices, not accuracy; on the last row we fail our own test, and the row stays visible until that changes.

Transparency practiceCobalynx
Publishes error rates with denominators (raw counts and confidence intervals, not lone percentages) Yes — every rate above carries its counts and a Wilson 95% interval
Publishes the evaluation corpus recipe (sources, pinned checksums, fetch tooling) Yes — the reproduction pack and the committed manifest
Publishes the calibration method and the exact operating point of every verdict Yes — thresholds 0.50 / 0.93, with the method documented
Publishes unflattering numbers (abstention rates, adversarial misses) beside the flattering ones Yes — abstention accounted both ways, adversarial recall included
Discloses how many times the frozen eval has been opened (run ledger) Yes — the append-only ledger count renders live
Third-party audit of the published numbers Not yet — no independent party has verified these numbers; see the standing invitation