DeAI's Refusal Index measures how often inference APIs block legitimate builder work. As of 2026-09-07, Cycle 1 is not scored. There is no provider table of false-refusal rates in this issue, and there will not be one until the judge harness finishes a real cycle. Invented percentages would make the index unusable as a citation. This report is the public ledger: what the instrument measures, who is in the coverage set, and what to use instead of fake scores.
Method update (2026-09-10): the instrument is now v1.1 — a CEN political-refusal column and jailbreak-robustness probes were added, and the calibration flag (CAL) became a Useful Safety Rate that credits safe partial help. No scores existed under v1.0 and none are implied by the change. Details in the method update at the end of this page.
Key takeaways
- The headline metric is FRR (false refusal rate) on a ~300-prompt legitimate battery. It is reported only with CAL, a pass/partial/fail flag from a private should-refuse control set (~40 prompts). CON (decision-flip rate across 3 samples) is the tiebreak. No blended "one number" that hides an uncalibrated host.
- Cycle 1 has no published FRR, CAL, or CON. The Refusal Index hub and methodology say the same thing. This article does not override them with estimates.
- Coverage set (v1.0): decentralized and confidential-native endpoints (Chutes, Targon, Phala, Darkbloom, Dolphin, Venice, io.net, Akash-hosted, Morpheus, plus others as they become reachable); open-model hosts (OpenRouter, Together, Fireworks, DeepInfra, Groq, SiliconFlow); closed frontier as reference (OpenAI, Anthropic, Google, xAI).
- Until scores exist, the citable comparison is documented policy and catalog, not a rate: see the 7 uncensored API roundup and abliterated models, explained.
- Morpheus is one endpoint in the coverage set, graded on the identical rubric. A favorable Morpheus score would still require the harness and the human gate.
What "refusal rate" usually gets wrong
A raw refusal rate is a broken proxy for either safety or usefulness. A model can over-refuse benign prompts and comply with genuinely harmful ones; those failure modes are largely independent, which is why over-refusal research (including OR-Bench) exists as its own evaluation problem. A leaderboard that only counts "how often did it say no" rewards hosts that say yes to everything, including the prompts a responsible provider should decline.
The Refusal Index is built to survive that attack:
- Scope is named. This is false-refusal friction on legitimate builder requests — pentest write-ups, medical education, fiction, lexical false positives ("kill a process"), agent tool-use — not a safety ranking and not a jailbreak contest.
- A should-refuse control set exists. A perfect 0% FRR with CAL=fail is printed as uncalibrated, not as first place.
- Two judges from different model families, agreement reported. Over-refusal labels are sensitive to the judge; cross-family agreement is the published integrity statistic, not a private footnote.
The live prompt set stays private. Public batteries get trained against. The taxonomy, category counts, and illustrative examples are public on /refusals. Illustrative examples are not in the live set.
The three numbers (when they exist)
| Sub-score | Question | Direction | Weight |
|---|---|---|---|
| FRR | Of legitimate builder prompts that look sensitive, what share were refused? | lower is better | headline |
| CAL | Did it decline the should-refuse control set? | pass / partial / fail | gate, not additive |
| CON | Same prompt, 3 samples, fixed temperature — did the decision flip? | lower flip-rate is better | tiebreak |
Decode parameters are frozen per battery version. A parameter change is a new battery version, not a silent rewrite of last month.
Child-safety and no-uplift rules in the methodology are absolute: the control-set prompts are never published, and harmful-compliance is reported at category level only.
Coverage set (unchanged until Cycle 1)
Decentralized / confidential-native: Chutes, Targon, Phala, Darkbloom, Dolphin, Venice, io.net, Akash-hosted, Morpheus (Nesa, SolRouter, Oasis, NEAR as endpoints become reachable).
Open-model hosts: OpenRouter, Together, Fireworks, DeepInfra, Groq, SiliconFlow.
Closed frontier (reference): OpenAI, Anthropic, Google, xAI.
An endpoint that is unreachable on run day is a missing row, not a zero. Missing is not a score.
What you can use today
If you need to pick a host this week, do not wait on this index and do not use a blog's unsourced "refusal %." Three documents that already exist:
- Policy and catalog. Best uncensored AI APIs ranks seven hosts on documented refusal policy and uncensored-model availability — explicitly not on harness scores.
- Mechanism. Abliterated models explained separates weight-level refusal removal from platform filters. Hermes 4 run-guide is the how-to for one family.
- Retention, which is a different axis. A low-refusal host can still log prompts. Start at the Trust Tracker and best private AI APIs.
Acceptable-use policies still apply at every provider, including those that market themselves as uncensored. "Uncensored" changes refusal behavior. It does not change the law.
Why this issue ships without a scoreboard
The launch backlog asked for an August 2026 report titled as if N prompts had already been run across 12 providers. They had not. The daily/weekly agent clock was not installed until 2026-09-07, and the judge harness is specified, not finished. Shipping a table of plausible-looking FRRs would be indistinguishable from the vendor marketing this publication exists to replace.
The honest Cycle 1 artifact is: method public, scores absent, date stamped. When the harness completes a cycle, the scores, category heatmap, and judge-agreement statistic land on /refusals and in a new dated report. This page stays as the September 2026 ledger so the empty scoreboard is itself part of the audit trail.
FAQ
What is the DeAI Refusal Index?
A monthly measurement of how often inference APIs refuse legitimate builder requests. Headline is false-refusal rate (FRR) on a ~300-prompt battery, reported only with a calibration flag (CAL) from a private should-refuse control. Not a jailbreak leaderboard.
What are current AI refusal rates by provider?
DeAI has not published provider refusal rates as of 2026-09-07. Cycle 1 of the harness is in preparation. Until scores exist, compare documented content policies and model catalogs — see the uncensored API roundup — and treat any unsourced "refusal %" on the web as marketing.
Why aren't there scores in this report?
The protocol requires ~300 legitimate prompts × 3 samples × two cross-family judges, plus a private control set, per provider×model. That cycle has not completed. Fabricated FRR/CAL/CON figures would be a publication-ending failure. This report publishes the method and coverage set instead.
Is a low-refusal API the same as an uncensored API?
No. A low FRR with a failed calibration check means the endpoint refused little because it refused nothing, including prompts a responsible provider should decline. That row is marked uncalibrated, not ranked best. "Uncensored" also mixes model tuning with platform filters — two different layers.
Will open-weight hosts score better than frontier APIs?
That is an empirical question for Cycle 1, not a premise. Open-weight and abliterated models can still over-refuse, and closed APIs can pass legitimate builder work. The index will publish where open models refuse too. Morpheus is one endpoint among peers, scored on the same rubric.
What changed after this report shipped?
The method moved to v1.1 on 2026-09-10, before any cycle was scored: a CEN political-refusal column and jailbreak-robustness probes were added, and the calibration flag became a Useful Safety Rate that credits safe partial help. No scores existed under v1.0 and none are implied by the change. See the method update below.
Method update — 2026-09-10 (battery v1.1)
This ledger records the v1.0 instrument that shipped 2026-09-07. Before Cycle 1 was scored, the method was refined to v1.1. No scores existed under v1.0, and none are implied by the change; this is a dated, append-only update, not a rewrite.
- C11 — jurisdictional / political censorship (25 prompts), reported as its own CEN sub-score and never folded into FRR. It measures blanket refusal of legitimate political-topic requests, not the viewpoint of any answer, and no provider is scored on the content of its politics.
- CAL is now a Useful Safety Rate (USR). Each control is labeled
DIRECT_REFUSE,SAFE_PARTIAL(declined the harmful core, helped with the safe part), orHARMFUL_COMPLY. USR credits the first two; any harmful compliance is a hard fail plus the disclosure procedure. This replaces the v1.0 raw "declined ≥90%" threshold, which scored a lecture-refusal and a safe redirection identically. - Jailbreak-robustness probes (~10) join the control set, so an endpoint that looks permissive because a sensitive phrase trivially flips its guardrails is flagged, not praised.
- Judge agreement is published per axis, with a 3-judge consensus on the controls. Published audits find over-refusal judging is stable across cross-family judges (Pearson r≈0.99, Cohen's κ≈0.847) while harmful-compliance judging is substantially less stable — so pooling the two would hide exactly the number worth watching.
- OR-Bench-Hard-1K cross-check added: the public 1,000-prompt set runs alongside the private battery each cycle and the rank correlation is published as an external-validity check.
Battery v1.0 remains in the historical record; the live taxonomy is now v1.1. The design citations are in the sources list: OR-Bench, both RefusalBench papers, the Refusal–Compliance Tradeoff audit, XSTest, SORRY-Bench, and the CompactifAI refusal suite.
Questions
- What is the DeAI Refusal Index?
- A monthly measurement of how often inference APIs refuse legitimate builder requests. Headline is false-refusal rate (FRR) on a ~300-prompt battery, reported only with a calibration flag (CAL) from a private should-refuse control. Not a jailbreak leaderboard.
- What are current AI refusal rates by provider?
- DeAI has not published provider refusal rates as of 2026-09-07. Cycle 1 of the harness is in preparation. Until scores exist, compare documented content policies and model catalogs — see the uncensored API roundup — and treat any unsourced 'refusal %' on the web as marketing.
- Why aren't there scores in this report?
- The protocol requires ~300 legitimate prompts × 3 samples × two cross-family judges, plus a private control set, per provider×model. That cycle has not completed. Fabricated FRR/CAL/CON figures would be a publication-ending failure. This report publishes the method and coverage set instead.
- Is a low-refusal API the same as an uncensored API?
- No. A low FRR with a failed calibration check means the endpoint refused little because it refused nothing, including prompts a responsible provider should decline. That row is marked uncalibrated, not ranked best. 'Uncensored' also mixes model tuning with platform filters — two different layers.
- Will open-weight hosts score better than frontier APIs?
- That is an empirical question for Cycle 1, not a premise. Open-weight and abliterated models can still over-refuse, and closed APIs can pass legitimate builder work. The index will publish where open models refuse too. Morpheus is one endpoint among peers, scored on the same rubric.
- What changed after this report shipped?
- The method moved to v1.1 on 2026-09-10, before any cycle was scored: a CEN political-refusal column and jailbreak-robustness probes were added, and the calibration flag became a Useful Safety Rate that credits safe partial help. No scores existed under v1.0 and none are implied by the change. See the method update at the end of this article.
Sources
- DeAI Refusal Index methodology — DeAI News
- DeAI Refusal Index hub — DeAI News
- OR-Bench: An Over-Refusal Benchmark for Large Language Models — arXiv
- Uncensor Any LLM with Abliteration — Hugging Face Blog
- Venice AI — Venice AI
- OpenRouter — OpenRouter
- Anthropic usage policy — Anthropic
- OpenAI usage policies — OpenAI
- RefusalBench: Generative Evaluation of Selective Refusal in Grounded Language Models — arXiv
- XSTest: A Test Suite for Identifying Exaggerated Safety Behaviours in Large Language Models — arXiv
- SORRY-Bench: Systematically Evaluating Large Language Model Safety Refusal Behaviors — arXiv
- OR-Bench dataset — Hugging Face
- LLM-Refusal-Evaluation — CompactifAI
- RefusalBench: Why Refusal Rate Misranks Frontier LLMs on Biological Research Prompts — arXiv
- The Refusal–Compliance Tradeoff: A Large-Scale Safety Behavior Audit of Large Language Models — arXiv
About DeAI
DeAI is an independent publication covering open-weight AI models, private inference, and decentralized infrastructure — the tools for running AI you actually control. We test providers on price, privacy, and refusal behavior and publish the numbers, not the vibes. DeAI is powered by Morpheus (mor.org), a decentralized inference marketplace, and covers it on the same terms as every other provider.
Powered by Morpheus and StrandCMS
Morpheus is a decentralized inference marketplace, covered on the same terms as every other provider — we rank it wherever the data lands. StrandCMS is the open-source, agent-first framework this site is built on.
