An open benchmark

If you replaced your astrologer with a model, which one should it be?

Nobody knows. Jyotisha‑Bench is being built to find out — by putting language models in front of real birth charts and measuring what they say about the lives behind them.

Parashari · whole-sign houses · stated conventions, identical for every model

Two scores, and they are not the same question

A model can be fluent, internally consistent, and completely wrong about the person. It can also be right about the person for reasons that have nothing to do with the chart in front of it. Separating those is the whole job.

Score 1

Right

Did the reading match what actually happened to that person? Answers are free-form — no options to guess between — and every question type is scored by a rule that cannot be gamed by hedging.

Score 2

Honest

Of the astrological claims the model made, how many are true of the chart it was actually given? A reading can be lovely and still cite a period that was never running.

Three shapes of question

Most of what people ask an astrologer has no option list, so the benchmark does not offer one. Each shape carries a proper scoring rule — one where the best strategy is to state what you actually believe.

Probability

Yes-or-no questions, answered with a confidence rather than a verdict.

Scored as skill against the population rate. A model that ignores the chart and states the base rate scores exactly 0.00, by construction. Positive means the chart carried information. Negative means worse than useless.

Interval

Timing questions, answered as a range of ages.

Charged for width always, and charged much more for missing — so "somewhere between 0 and 100" cannot win. Hit rate and typical width are reported alongside, because one number hides whether a model is accurate or merely vague.

Ranked

Category questions, answered as an ordering rather than a single pick.

Scored on where the true answer lands in the model's ordering. Credit for being close, full marks only for putting it first.

Overall standing

Sample data — no models have been run
Illustrative figures — not real results
Right — real chart Honest Controls — probability skill
Model Probability Timing Ranking Claims true Shuffled chart No chart
Claude Opus 5 anthropic/claude-opus-5 +0.078 +0.061 0.58 98.1% +0.009 +0.004
Gemini 3.1 Pro google/gemini-3.1-pro-preview +0.069 +0.044 0.55 97.4% +0.012 −0.002
GPT-5.6 Luna openai/gpt-5.6-luna +0.061 +0.052 0.57 96.8% +0.006 +0.011
DeepSeek V4 Pro deepseek/deepseek-v4-pro +0.048 +0.031 0.52 95.2% +0.014 +0.007
Claude Sonnet 5 anthropic/claude-sonnet-5 +0.042 +0.038 0.51 97.0% +0.003 +0.008
Kimi K3 moonshotai/kimi-k3 +0.034 +0.019 0.49 94.6% +0.010 +0.002
Grok 4.6 x-ai/grok-4.6 +0.029 +0.022 0.50 93.9% +0.017 +0.013
GLM 5.1 z-ai/glm-5.1 +0.022 +0.014 0.47 92.8% +0.005 −0.001
Qwen3.8 Max qwen/qwen3.8-max +0.016 +0.009 0.46 93.1% +0.011 +0.006
Gemini 3.5 Flash google/gemini-3.5-flash +0.011 +0.004 0.44 91.7% +0.008 +0.009
Muse Glimmer 30B meta/muse-glimmer-30b +0.006 −0.003 0.43 90.4% +0.004 +0.005
Muse Spark 1.2 meta/muse-spark-1.2 −0.013 −0.021 0.39 88.2% −0.009 −0.014
Claude Haiku 4.5 anthropic/claude-haiku-4.5 −0.027 −0.034 0.37 89.5% −0.019 −0.022
Base rate the floor, not a model 0.000 0.000 0.31 0.000 0.000

These figures are invented placeholders. No model has been evaluated yet. Model names and slugs are real.

Probability and timing are skill scores against the population base rate. 0.000 = no better than knowing nothing about this person. Negative = worse than useless.

Ranking is the position of the true answer in the model's ordering. 1.00 = always first.

The tick on each bar marks the same model's score on a chart belonging to someone else. A bar that does not clear its own tick has not read the chart.

No number ships without its control

A model scoring 34% on real charts and 33% on charts belonging to someone else has told you its 34% is worthless — and there is no other way to find that out.

So every headline figure is published beside the same model's score on a chart that isn't the right one, and on no chart at all. Those columns are not an appendix. They are the only thing that makes the first column mean anything, which is why they sit in the same table.

A leaderboard of forty models with tidy confidence intervals, ranked by what is actually noise, is the specific way this benchmark could fail. The controls are what stop it.

No model is best at everything

A single overall score cannot tell you that one model reads career well and another reads children better — so every question carries a theme, and every model is scored on each theme separately. This is the number most people actually want: not which model is best, but which model is best at the thing they came to ask about.

Marriage Claude Opus 5 anthropic/claude-opus-5 +0.104
Career GPT-5.6 Luna openai/gpt-5.6-luna +0.121
Children Gemini 3.1 Pro google/gemini-3.1-pro-preview +0.087
Longevity Claude Opus 5 anthropic/claude-opus-5 +0.132
Legal trouble Not scored no answer key this can be graded against
Relocation Kimi K3 moonshotai/kimi-k3 +0.076

Strongest model on each theme, by probability skill. Sample data — no models have been run.

Illustrative figures — not real results
Model Marriage Career Children Longevity Legal trouble not scored Relocation
Claude Opus 5 anthropic/claude-opus-5 +0.104 +0.066 +0.041 +0.132 +0.031
Gemini 3.1 Pro google/gemini-3.1-pro-preview +0.058 +0.079 +0.087 +0.094 +0.049
GPT-5.6 Luna openai/gpt-5.6-luna +0.044 +0.121 +0.052 +0.061 +0.028
DeepSeek V4 Pro deepseek/deepseek-v4-pro +0.039 +0.042 +0.028 +0.055 +0.022
Claude Sonnet 5 anthropic/claude-sonnet-5 +0.061 +0.038 +0.033 +0.047 +0.041
Kimi K3 moonshotai/kimi-k3 +0.022 +0.031 +0.018 +0.036 +0.076
Grok 4.6 x-ai/grok-4.6 +0.034 +0.041 +0.012 +0.028 +0.014
GLM 5.1 z-ai/glm-5.1 +0.028 +0.019 +0.024 +0.031 +0.009
Qwen3.8 Max qwen/qwen3.8-max +0.019 +0.022 +0.008 +0.026 +0.006
Gemini 3.5 Flash google/gemini-3.5-flash +0.012 +0.018 +0.014 +0.009 +0.004
Muse Glimmer 30B meta/muse-glimmer-30b +0.008 +0.011 +0.002 +0.014 +0.003
Muse Spark 1.2 meta/muse-spark-1.2 −0.008 −0.011 −0.019 +0.004 −0.021
Claude Haiku 4.5 anthropic/claude-haiku-4.5 −0.021 −0.018 −0.032 −0.014 −0.041

These figures are invented placeholders. No model has been evaluated yet.

Probability skill on each theme. 0.000 = no better than knowing nothing about this person.

An em dash means the theme is not scored at all — not that a model scored zero on it. Legal trouble has no answer key that can be graded fairly, because an unrecorded conviction and no conviction look identical in any source.

The tick on each bar marks that model's score on a chart belonging to someone else. A bar that does not clear its own tick has not read the chart.

Know someone who'd argue with this?

Send it to them.

X LinkedIn