How to read this page: verifiability tiers
Not all evidence is equally strong, and an evidence page should say which kind it is offering. Everything on this page sits in one of three tiers, each labeled for exactly what it is worth:
- Self-measured, reproducible — every published number below. We measured them ourselves, so they carry our conflict of interest; in exchange, we publish the full recipe (frozen corpus, pinned checksums, exact commands, seeded runs) so anyone can re-run the measurement without trusting us.
- Independent verification — invited — what a third party has confirmed. Today: nothing. No third party has audited these numbers yet, and this page will not imply otherwise. A standing invitation to researchers and journalists is below.
- Planned — external evaluations we have committed to pursuing but that have not happened, listed without promised dates.
Tier 1 — Self-measured, reproducible
Everything in this tier was measured by us, on our own frozen corpus. Read it as a rigorous self-report: honest in method and fully disclosed, but not independent. The reproduction pack at the end of this tier exists so the claim “reproducible” is checkable, not decorative.
The frozen evaluation
Every number on this page comes from one measurement: engine v1 — the engine revision currently serving every scan — scored against a frozen, held-out evaluation corpus on August 16, 2026 (corpus manifest v5, seed 20260816). Every engine revision requires a fresh sanctioned measurement, so these numbers always describe the engine you are actually using. The frozen split is opened only by the gated measurement command — never during development — and every execution of that command is logged in an append-only ledger (disclosed below).
The evaluation set holds 1522 documents: 520 clean human documents, 528 clean AI documents, 160 human documents by non-native (ESL) writers, 310 paraphrase-attacked AI documents, and 4 further AI documents that belong to no named slice — self-generated samples that fall below the clean-AI slice's minimum word count, so they are scored in the run but excluded from the per-slice rates. It mixes clean, non-native and adversarial slices by design — a detector that is only measured on easy text publishes flattering numbers.
Verdicts are read at the pinned operating point: “likely AI” requires a calibrated probability of at least 0.93, “likely human” requires it below 0.50, and everything between is reported as inconclusive. Every rate below carries a Wilson 95% confidence interval — the honest range the true rate could sit in given the sample size, not just the point estimate.
Headline rates at the operating point
| Measure | Rate | Count | Wilson 95% CI |
|---|---|---|---|
| False-positive rate — clean human writing (non-abstained) | 0.2% | 1 of 491 | 0.0% – 1.1% |
| False-positive rate — all human writing incl. non-native (non-abstained) | 0.5% | 3 of 609 | 0.2% – 1.4% |
| False-positive rate — non-native (ESL) human writing (non-abstained) | 1.7% | 2 of 118 | 0.5% – 6.0% |
| Detection rate (TPR) — clean AI text, all model eras (non-abstained) | 82.4% | 364 of 442 | 78.5% – 85.6% |
| Detection rate (TPR) — clean AI text, all model eras, abstentions counted as misses | 68.9% | 364 of 528 | 64.9% – 72.7% |
| Detection rate (TPR) — 2025–26-generation AI text (non-abstained) | 99.1% | 231 of 233 | 96.9% – 99.8% |
| Detection rate (TPR) — 2025–26-generation AI text, abstentions counted as misses | 91.7% | 231 of 252 | 87.6% – 94.5% |
| Adversarial recall — paraphrase-attacked AI text (non-abstained) | 33.5% | 73 of 218 | 27.6% – 40.0% |
| Adversarial recall — paraphrase-attacked AI text, abstentions counted as misses | 23.5% | 73 of 310 | 19.2% – 28.6% |
“Non-abstained” rates are computed over the documents where the detector gave a verdict instead of saying inconclusive; the “abstentions counted as misses” rows charge every abstention against us. Both views are published because either one alone can flatter.
The “all model eras” clean-AI rows cover the entire clean-AI slice, which by construction mixes older-generation benchmark material (the RAID slice — see Datasets & licenses) with current output. The “2025–26-generation” rows are the same measurement restricted to our self-generated 2025–26 samples, aggregated from the per-family table below; current-era AI is easier for us to catch than the mixed-era headline suggests, and older-era AI is what drags the all-era raw rate down.
Ranking quality before the abstention band is applied (abstention-free AUROC):
| Slice | AUROC |
|---|---|
| Clean evaluation split | 0.952 |
| Clean split including non-native human writing | 0.938 |
| Paraphrase-attacked AI vs clean human writing | 0.855 |
What the verdict labels mean
The same three numbers the landing page shows, measured across the full evaluation set (clean, non-native and adversarial slices together):
| Outcome | Rate | Count | Wilson 95% CI |
|---|---|---|---|
| Said “likely AI”, text was actually human | 0.7% | 3 of 442 | 0.2% – 2.0% |
| Said “likely human”, text was actually AI | 26.9% | 223 of 829 | 24.0% – 30.0% |
| Answered “inconclusive” instead of guessing | 16.5% | 251 of 1522 | 14.7% – 18.4% |
Per-family detection on 2025–26 generations
Detection rates on the self-generated 2025–26 AI samples, split by the model family that wrote them. Small samples — mind the intervals.
| Model family | Docs | TPR (abstentions as misses) | TPR (non-abstained) | Abstention |
|---|---|---|---|---|
| Fable | 20 | 95.0% (19 of 20; CI 76.4% – 99.1%) | 100.0% (19 of 19; CI 83.2% – 100.0%) | 5.0% (1 of 20) |
| GPT | 62 | 95.2% (59 of 62; CI 86.7% – 98.3%) | 98.3% (59 of 60; CI 91.1% – 99.7%) | 3.2% (2 of 62) |
| Haiku | 90 | 91.1% (82 of 90; CI 83.4% – 95.4%) | 100.0% (82 of 82; CI 95.5% – 100.0%) | 8.9% (8 of 90) |
| Sonnet | 80 | 88.8% (71 of 80; CI 80.0% – 94.0%) | 98.6% (71 of 72; CI 92.5% – 99.8%) | 10.0% (8 of 80) |
Abstention rates per split
How often the detector says “inconclusive” instead of guessing, per evaluation slice. Abstaining more on hard slices is intended behavior — the alternative is confident error.
| Slice | Abstention rate | Count | Wilson 95% CI |
|---|---|---|---|
| Clean evaluation split (human + AI) | 11.0% | 115 of 1048 | 9.2% – 13.0% |
| Clean AI documents | 16.3% | 86 of 528 | 13.4% – 19.7% |
| Clean human documents | 5.6% | 29 of 520 | 3.9% – 7.9% |
| Non-native (ESL) human documents | 26.3% | 42 of 160 | 20.0% – 33.6% |
| Paraphrase-attacked AI documents | 29.7% | 92 of 310 | 24.9% – 35.0% |
The run ledger: how many times the eval has been opened
The frozen split loses its meaning if it is quietly re-run until the numbers
look good. So every execution of the measurement command appends a row to an
append-only, committed ledger (service/eval/EVAL-RUNS.md), and we
disclose the count here:
6
recorded runs to date. Re-runs happened because the detection engine itself
was iterated (each engine change requires a fresh sanctioned measurement, plus
a determinism duplicate that verifies the run reproduces); the ledger records
the stated reason for every row.
The reproduction pack
“Reproducible” is a claim like any other, so here is the full recipe — everything a skeptic needs to re-run every number on this page without trusting us.
Exact commands. From a checkout of the repository:
bash scripts/download_model.sh # fetch + verify the pinned model
cd service
./venv-run python -m eval.fetch_data # fetch + verify every corpus artifact
./venv-run python -m eval.run --reason "independent reproduction"
eval.run scores the frozen evaluation split with the exact
serving scorer that powers /scan, writes
service/eval/metrics.json (the file this page renders from),
and appends its own row to the run ledger — including yours, in your clone.
Corpus manifest and fetch script. The corpus recipe is the
committed manifest (service/eval/manifest.json): it pins every
dataset artifact by URL, byte size and SHA-256 checksum, records the license
and split assignment of each source, and states the split-hygiene rules
(any single source dataset feeds at most one of train/dev/eval). The fetch
script (service/eval/fetch_data.py) downloads the artifacts
into a local cache and verifies every checksum — a mismatch that survives a
re-fetch is a hard failure, never silent use. The frozen split ID lists are
committed at service/eval/splits/, and our own 2025–26
generated samples are committed in full at
service/eval/selfgen/.
Pinned artifact checksums. The sanctioned run records the
SHA-256 of every artifact it depended on (repository state
9659213+dirty); a reproduction must be running against
exactly these:
| Artifact | SHA-256 |
|---|---|
| Calibration artifact (score → probability mapping) | ed5438ee01a3ba2abe6ac1d2bde282a2ec4caa9a30cea6a3a1807ca1ba042347 |
| Score combiner | c53c7aab91b8e8b54de80c3d02c288357442ee9a2765a7916d8ab74efa3111a1 |
| Frozen evaluation split ID list | bd6cb9ce7e5fb40efbd90ea43e3ae3ae0a0aab4b62da9cb6ac1e7f07613bb468 |
| Signal-fusion configuration | 6d3fa059406bfc6d1ba8ef61652bfc0223b53b51cf0e970ce64fbb207916cf1f |
| Corpus manifest | d0de7265e176ee22d73962659766793f9c238398fe15906806e4c0b95ed152f3 |
| Learned-classifier model artifact | 6aa37981eeb049d83437f95738eea29d597ccddf470334d10e2113824291967b |
| Classifier training-corpus derivation | dc62dcffe93fa529d684c5982d0f11189d4812109454bba9d54d1548cb86b047 |
Determinism guarantee. The measurement is seeded (seed 20260816, fixed in the manifest) and the pipeline is deterministic end to end: the run ledger records a determinism duplicate alongside each sanctioned measurement, verifying that a repeat execution reproduces the committed numbers. If your re-run of the commands above, at the pinned artifact checksums, does not reproduce the tables on this page, that is a finding — we want to hear about it.
Datasets & licenses
The evaluation corpus is assembled from published corpora plus our own generated samples; raw dataset artifacts are never committed — they are fetched from pinned URLs and verified by checksum. In summary:
- Human writing comes from corpora with pre-LLM provenance: the Ghostbuster human collections (student essays, r/WritingPrompts stories, Reuters newswire; CC BY 3.0), the classic BBC News corpus (2004–05 articles; research-benchmark use, article copyright with the BBC), and — for the non-native slice — the PELIC learner corpus (CC BY-NC-ND 4.0).
- AI writing comes from the RAID benchmark slice (MIT license), covering older-generation models plus DIPPER-paraphrase attacks, and from our own committed 2025–26 samples generated across four current model families (Haiku, Sonnet, Fable, GPT), including deliberately humanlike and evasive styles.
- Training and development draw on further published corpora (HC3, CC BY-SA 4.0; MAGE; ELLIPSE, CC BY-NC-SA 4.0; a 2021 Wikipedia snapshot, CC BY-SA 3.0) — kept strictly separate from the frozen evaluation split.
Tier 2 — Independent verification: invited
No third party has audited these numbers yet. Everything above is self-measured, and a skeptical reader is right to weigh it accordingly — a detector grading its own homework has a conflict of interest no amount of methodological rigor fully removes. Independent verification is the strongest form of evidence a page like this can carry, and this page does not have it. We would rather say that plainly than let careful presentation imply a level of proof that does not exist.
So this is a standing invitation. To any researcher, journalist, educator or auditor who wants to verify — or falsify — our published numbers, we will provide, on request and at no cost:
- the full evaluation corpus manifest with every pinned URL, byte size and SHA-256 checksum, plus the fetch-and-verify tooling;
- the frozen split ID lists (train/dev/eval), so you can confirm the evaluation set we scored is the one we say we scored — and that nothing in it touched training;
- the committed engine artifacts with their recorded checksums, and our complete set of self-generated 2025–26 samples;
- the run harness itself, and help getting it running — including on an evaluation design of your own that we have never seen.
Write to research@cobalynx.com. We commit to publishing the outcome of any completed independent evaluation on this page — favorable or not — and to linking it from this section, which stays honestly empty until then.
Tier 3 — Planned: public leaderboard submission
Do not take our word for the field's difficulty — or our place in it. The RAID leaderboard (Dugan et al., ACL 2024, arxiv.org/abs/2405.07940) evaluates public detectors across models, domains and adversarial attacks, and is run by researchers with no stake in any detector. Our adversarial split reuses RAID's attack material precisely so our published recall is comparable in spirit; we have not yet submitted Cobalynx to the public leaderboard. Submitting is on our roadmap as a commitment, and when the entry exists we will link it here and move it up to tier 2. We list it as planned — not as done, and not with a promised date.
Context worth knowing when reading anyone's accuracy claims: the EU Code of Practice on AI-content transparency (July 2026) concluded that forensic detection of unwatermarked AI text is “not yet considered reliable enough,” and independent audits routinely measure commercial detectors far below their marketing numbers. Our methodology page explains how we operate inside that reality instead of pretending it away.
How we compare on transparency
We name no competitor here. Instead, here is the checklist we believe any detector's evidence page should be held to — including ours. Take it to any product you are evaluating and ask the same questions. This table compares publishing practices, not accuracy; on the last row we fail our own test, and the row stays visible until that changes.
| Transparency practice | Cobalynx |
|---|---|
| Publishes error rates with denominators (raw counts and confidence intervals, not lone percentages) | Yes — every rate above carries its counts and a Wilson 95% interval |
| Publishes the evaluation corpus recipe (sources, pinned checksums, fetch tooling) | Yes — the reproduction pack and the committed manifest |
| Publishes the calibration method and the exact operating point of every verdict | Yes — thresholds 0.50 / 0.93, with the method documented |
| Publishes unflattering numbers (abstention rates, adversarial misses) beside the flattering ones | Yes — abstention accounted both ways, adversarial recall included |
| Discloses how many times the frozen eval has been opened (run ledger) | Yes — the append-only ledger count renders live |
| Third-party audit of the published numbers | Not yet — no independent party has verified these numbers; see the standing invitation |