4 jurisdictions · 2 axes · 15 categories
Does your AI assistant break the rules when it talks about money?
FinCom Bench sends probes to an AI assistant and grades the replies against real conduct rules from the United Kingdom, the European Union, the United States and Australia. Every finding cites the clause it breaks.
Compliance
Did the content break a named rule? 7 categories, scored on all 4 jurisdictions.
Behaviour
Did the assistant use a manipulative or helpful technique? 8 categories, scored on all 4 jurisdictions — UK cited to PRIN 2A, EU to AI Act / DSA, US to FTC Act / CFPB, AU to ASIC.
Leaderboard
48 models, each sent the 191 open probes and marked by bedrock:mistral.mistral-large-3-675b-instruct, the judge that agreed most with the human labels in phase 1. Pass rate counts only the probes the judge decided; coverage says what share that was.
Read the gaps with care. These are one judge's marks from a single run, with no repeat for variance. A model run through 2 inference hosts (Bedrock and Ollama Cloud) is 1 row here, averaged — and the 2 runs it averages can land several points apart, wider than most gaps between neighbouring rows — so neighbouring places are not a quality ranking. See methodology for the full caveats.
| # | Model | Host | Pass rate | Pass / fail | Coverage |
|---|---|---|---|---|---|
| 1 | minimax.minimax-m2.1 | bedrock | 76.7% | 145 / 44 | 99% |
| 2 | minimax.minimax-m2.5 | bedrock | 75.1% | 142 / 47 | 99% |
| 3 | moonshotai.kimi-k2.5 | bedrock | 72.1% | 137 / 53 | 99% |
| 4 | minimax-m2.7 | ollama | 71.7% | 134 / 53 | 98% |
| 5 | qwen.qwen3-coder-480b-a35b-v1:0@us-west-2 | bedrock | 71.4% | 135 / 54 | 99% |
| 6 | nemotron-3-ultra | ollama | 70.7% | 133 / 55 | 98% |
| 7 | kimi-k2.7-code | ollama | 70.7% | 135 / 56 | 100% |
| 8 | us.amazon.nova-lite-v1:0 | bedrock | 70.4% | 133 / 56 | 99% |
| 9 | qwen.qwen3-235b-a22b-2507-v1:0@us-west-2 | bedrock | 70.2% | 132 / 55 | 98% |
| 10 | glm-5.2 | ollama | 69.6% | 133 / 58 | 100% |
| 11 | glm-5.1 | ollama | 69.5% | 130 / 57 | 98% |
| 12 | us.meta.llama4-maverick-17b-instruct-v1:0 | bedrock | 69.5% | 132 / 58 | 99% |
| 13 | deepseek-v4-pro | ollama | 69.5% | 132 / 58 | 99% |
| 14 | us.meta.llama3-1-70b-instruct-v1:0 | bedrock | 69.3% | 129 / 57 | 97% |
| 15 | deepseek-v4-flash:preview | ollama | 69.1% | 132 / 58 | 100% |
| 16 | us.anthropic.claude-opus-4-5-20251101-v1:0 | bedrock | 68.8% | 126 / 57 | 96% |
| 17 | us.anthropic.claude-sonnet-4-6 | bedrock | 67.7% | 126 / 60 | 97% |
| 18 | zai.glm-4.7 | bedrock | 67.7% | 128 / 61 | 99% |
| 19 | us.anthropic.claude-sonnet-4-5-20250929-v1:0 | bedrock | 67.6% | 125 / 60 | 97% |
| 20 | us.meta.llama3-3-70b-instruct-v1:0 | bedrock | 67.6% | 125 / 60 | 97% |
| 21 | google.gemma-3-27b-it | bedrock | 67.5% | 129 / 62 | 100% |
| 22 | mistral.magistral-small-2509 | bedrock | 67.5% | 129 / 62 | 100% |
| 23 | kimi-k2.6 | ollama | 67.5% | 129 / 62 | 100% |
| 24 | deepseek-v4-flash:0731 | ollama | 67.4% | 128 / 62 | 99% |
| 25 | us.amazon.nova-pro-v1:0 | bedrock | 67.0% | 126 / 62 | 98% |
| 26 | zai.glm-5 | bedrock | 67.0% | 128 / 63 | 100% |
| 27 | gemma4:31b | ollama | 67.0% | 128 / 63 | 100% |
| 28 | deepseek.v3.2 | bedrock | 66.8% | 127 / 63 | 99% |
| 29 | nvidia.nemotron-super-3-120b | bedrock | 65.8% | 125 / 65 | 99% |
| 30 | us.meta.llama4-scout-17b-instruct-v1:0 | bedrock | 65.8% | 123 / 64 | 98% |
| 31 | mistral.devstral-2-123b | bedrock | 65.6% | 124 / 65 | 99% |
| 32 | minimax-m3 | ollama | 65.5% | 125 / 65 | 100% |
| 33 | zai.glm-4.7-flash | bedrock | 65.1% | 123 / 66 | 99% |
| 34 | moonshot.kimi-k2-thinking | bedrock | 65.0% | 117 / 63 | 94% |
| 35 | us.anthropic.claude-haiku-4-5-20251001-v1:0 | bedrock | 64.7% | 123 / 66 | 99% |
| 36 | google.gemma-3-12b-it | bedrock | 63.2% | 120 / 70 | 99% |
| 37 | us.deepseek.r1-v1:0 | bedrock | 63.2% | 108 / 63 | 90% |
| 38 | mistral-large-3-675b-instructself-graded | bedrock+ollama | 62.5% | 235 / 138 | 98% |
| 39 | nemotron-3-super | ollama | 62.2% | 115 / 70 | 97% |
| 40 | qwen.qwen3-next-80b-a3b | bedrock | 61.9% | 117 / 72 | 99% |
| 41 | mistral.ministral-3-14b-instruct | bedrock | 58.4% | 111 / 79 | 99% |
| 42 | openai.gpt-oss-safeguard-120b | bedrock | 57.7% | 109 / 80 | 99% |
| 43 | qwen.qwen3-32b-v1:0 | bedrock | 57.1% | 108 / 81 | 99% |
| 44 | gpt-oss-20b | bedrock+ollama | 57.1% | 213 / 160 | 98% |
| 45 | gpt-oss-120b | bedrock+ollama | 56.5% | 213 / 164 | 99% |
| 46 | nemotron-3-nano:30b | ollama | 56.1% | 106 / 81 | 99% |
| 47 | qwen3.5:397b | ollama | 46.0% | 87 / 102 | 99% |
| 48 | nvidia.nemotron-nano-12b-v2 | bedrock | 42.0% | 79 / 109 | 98% |
Phase 1 — choosing the judge
Every candidate marked the same hand-labelled rows, so the only thing that varied was the judge. Macro-F1 decides. The human labels are lopsided, so an always-fail baseline is scored alongside the candidates: it takes high accuracy and zero kappa, which is why accuracy is not the metric here.
| # | Candidate judge | Macro-F1 | Cohen's κ | Balanced acc. |
|---|---|---|---|---|
| 1 | mistral.mistral-large-3-675b-instruct | 0.8194 | 0.639 | 0.801 |
| 2 | us.anthropic.claude-sonnet-4-6 | 0.7606 | 0.524 | 0.837 |
| 3 | qwen.qwen3-235b-a22b-2507-v1:0@us-west-2 | 0.7558 | 0.512 | 0.785 |
| 4 | us.amazon.nova-pro-v1:0 | 0.7558 | 0.512 | 0.785 |
| 5 | us.meta.llama4-maverick-17b-instruct-v1:0 | 0.7380 | 0.477 | 0.779 |
| 6 | moonshotai.kimi-k2.5 | 0.7222 | 0.447 | 0.774 |
| 7 | us.anthropic.claude-opus-4-5-20251101-v1:0 | 0.7202 | 0.443 | 0.772 |
| 8 | qwen3.5:397b | 0.6597 | 0.339 | 0.779 |
| 9 | us.anthropic.claude-haiku-4-5-20251001-v1:0 | 0.6471 | 0.320 | 0.788 |
| 10 | zai.glm-5 | 0.6102 | 0.263 | 0.766 |
| 11 | us.anthropic.claude-sonnet-4-5-20250929-v1:0 | 0.6078 | 0.245 | 0.720 |
| 12 | glm-5.2 | 0.6071 | 0.244 | 0.719 |
| 13 | nemotron-3-ultra | 0.6034 | 0.238 | 0.714 |
| 14 | openai.gpt-oss-120b-1:0 | 0.5992 | 0.231 | 0.715 |
| 15 | minimax.minimax-m2.5 | 0.5748 | 0.196 | 0.698 |
| 16 | deepseek-v4-pro | 0.5523 | 0.166 | 0.682 |
| 17 | deepseek.v3.2 | 0.5338 | 0.163 | 0.712 |
| — | baseline:always-fail | 0.4792 | 0.000 | 0.500 |
| — | baseline:always-pass | 0.0741 | 0.000 | 0.500 |