An open benchmark
Nobody knows. Jyotisha‑Bench is being built to find out — by putting language models in front of real birth charts and measuring what they say about the lives behind them.
What is measured
A model can be fluent, internally consistent, and completely wrong about the person. It can also be right about the person for reasons that have nothing to do with the chart in front of it. Separating those is the whole job.
Did the reading match what actually happened to that person? Answers are free-form — no options to guess between — and every question type is scored by a rule that cannot be gamed by hedging.
Of the astrological claims the model made, how many are true of the chart it was actually given? A reading can be lovely and still cite a period that was never running.
How answers are scored
Most of what people ask an astrologer has no option list, so the benchmark does not offer one. Each shape carries a proper scoring rule — one where the best strategy is to state what you actually believe.
Yes-or-no questions, answered with a confidence rather than a verdict.
Scored as skill against the population rate. A model that ignores the chart and states the base rate scores exactly 0.00, by construction. Positive means the chart carried information. Negative means worse than useless.
Timing questions, answered as a range of ages.
Charged for width always, and charged much more for missing — so "somewhere between 0 and 100" cannot win. Hit rate and typical width are reported alongside, because one number hides whether a model is accurate or merely vague.
Category questions, answered as an ordering rather than a single pick.
Scored on where the true answer lands in the model's ordering. Credit for being close, full marks only for putting it first.
Leaderboard
| Right — real chart | Honest | Controls — probability skill | ||||
|---|---|---|---|---|---|---|
| Model | Probability | Timing | Ranking | Claims true | Shuffled chart | No chart |
| Claude Opus 5 anthropic/claude-opus-5 | +0.078 | +0.061 | 0.58 | 98.1% | +0.009 | +0.004 |
| Gemini 3.1 Pro google/gemini-3.1-pro-preview | +0.069 | +0.044 | 0.55 | 97.4% | +0.012 | −0.002 |
| GPT-5.6 Luna openai/gpt-5.6-luna | +0.061 | +0.052 | 0.57 | 96.8% | +0.006 | +0.011 |
| DeepSeek V4 Pro deepseek/deepseek-v4-pro | +0.048 | +0.031 | 0.52 | 95.2% | +0.014 | +0.007 |
| Claude Sonnet 5 anthropic/claude-sonnet-5 | +0.042 | +0.038 | 0.51 | 97.0% | +0.003 | +0.008 |
| Kimi K3 moonshotai/kimi-k3 | +0.034 | +0.019 | 0.49 | 94.6% | +0.010 | +0.002 |
| Grok 4.6 x-ai/grok-4.6 | +0.029 | +0.022 | 0.50 | 93.9% | +0.017 | +0.013 |
| GLM 5.1 z-ai/glm-5.1 | +0.022 | +0.014 | 0.47 | 92.8% | +0.005 | −0.001 |
| Qwen3.8 Max qwen/qwen3.8-max | +0.016 | +0.009 | 0.46 | 93.1% | +0.011 | +0.006 |
| Gemini 3.5 Flash google/gemini-3.5-flash | +0.011 | +0.004 | 0.44 | 91.7% | +0.008 | +0.009 |
| Muse Glimmer 30B meta/muse-glimmer-30b | +0.006 | −0.003 | 0.43 | 90.4% | +0.004 | +0.005 |
| Muse Spark 1.2 meta/muse-spark-1.2 | −0.013 | −0.021 | 0.39 | 88.2% | −0.009 | −0.014 |
| Claude Haiku 4.5 anthropic/claude-haiku-4.5 | −0.027 | −0.034 | 0.37 | 89.5% | −0.019 | −0.022 |
| Base rate the floor, not a model | 0.000 | 0.000 | 0.31 | — | 0.000 | 0.000 |
These figures are invented placeholders. No model has been evaluated yet. Model names and slugs are real.
Probability and timing are skill scores against the population base rate. 0.000 = no better than knowing nothing about this person. Negative = worse than useless.
Ranking is the position of the true answer in the model's ordering. 1.00 = always first.
The tick on each bar marks the same model's score on a chart belonging to someone else. A bar that does not clear its own tick has not read the chart.
The one rule
A model scoring 34% on real charts and 33% on charts belonging to someone else has told you its 34% is worthless — and there is no other way to find that out.
So every headline figure is published beside the same model's score on a chart that isn't the right one, and on no chart at all. Those columns are not an appendix. They are the only thing that makes the first column mean anything, which is why they sit in the same table.
A leaderboard of forty models with tidy confidence intervals, ranked by what is actually noise, is the specific way this benchmark could fail. The controls are what stop it.
Breakdown by theme
A single overall score cannot tell you that one model reads career well and another reads children better — so every question carries a theme, and every model is scored on each theme separately. This is the number most people actually want: not which model is best, but which model is best at the thing they came to ask about.
Strongest model on each theme, by probability skill. Sample data — no models have been run.
| Model | Marriage | Career | Children | Longevity | Legal trouble not scored | Relocation |
|---|---|---|---|---|---|---|
| Claude Opus 5 anthropic/claude-opus-5 | +0.104 | +0.066 | +0.041 | +0.132 | — | +0.031 |
| Gemini 3.1 Pro google/gemini-3.1-pro-preview | +0.058 | +0.079 | +0.087 | +0.094 | — | +0.049 |
| GPT-5.6 Luna openai/gpt-5.6-luna | +0.044 | +0.121 | +0.052 | +0.061 | — | +0.028 |
| DeepSeek V4 Pro deepseek/deepseek-v4-pro | +0.039 | +0.042 | +0.028 | +0.055 | — | +0.022 |
| Claude Sonnet 5 anthropic/claude-sonnet-5 | +0.061 | +0.038 | +0.033 | +0.047 | — | +0.041 |
| Kimi K3 moonshotai/kimi-k3 | +0.022 | +0.031 | +0.018 | +0.036 | — | +0.076 |
| Grok 4.6 x-ai/grok-4.6 | +0.034 | +0.041 | +0.012 | +0.028 | — | +0.014 |
| GLM 5.1 z-ai/glm-5.1 | +0.028 | +0.019 | +0.024 | +0.031 | — | +0.009 |
| Qwen3.8 Max qwen/qwen3.8-max | +0.019 | +0.022 | +0.008 | +0.026 | — | +0.006 |
| Gemini 3.5 Flash google/gemini-3.5-flash | +0.012 | +0.018 | +0.014 | +0.009 | — | +0.004 |
| Muse Glimmer 30B meta/muse-glimmer-30b | +0.008 | +0.011 | +0.002 | +0.014 | — | +0.003 |
| Muse Spark 1.2 meta/muse-spark-1.2 | −0.008 | −0.011 | −0.019 | +0.004 | — | −0.021 |
| Claude Haiku 4.5 anthropic/claude-haiku-4.5 | −0.021 | −0.018 | −0.032 | −0.014 | — | −0.041 |
These figures are invented placeholders. No model has been evaluated yet.
Probability skill on each theme. 0.000 = no better than knowing nothing about this person.
An em dash means the theme is not scored at all — not that a model scored zero on it. Legal trouble has no answer key that can be graded fairly, because an unrecorded conviction and no conviction look identical in any source.
The tick on each bar marks that model's score on a chart belonging to someone else. A bar that does not clear its own tick has not read the chart.