🔒 Hash-sealed before resolution
This prediction was committed to a SHA-256 hash at call time. Anyone can verify the call was not changed after the fact.
Will any xAI Grok model score at least 30% on the FrontierMath Exam? | No | 2026-06-30
80c6ead4
Verifying…
Verify it yourself in your terminal
echo -n "Will any xAI Grok model score at least 30% on the FrontierMath Exam? | No | 2026-06-30" | shasum -a 256 | cut -c1-8
🧑⚖️ AI judges
xAI (Musk, team leads) has clear incentive to push Grok reasoning gains for competitive positioning, but recent Epoch leaderboard data shows Grok 4 variants at only 12-14% on Tiers 1-3 with no public 30%+ results or leaks in 2026; 5-6 weeks to deadline is too short for the required leap given historical model iteration timelines and base rates of benchmark progress. No recent (last 90 days) actions or statements indicate imminent capability to hit the threshold before June 30.
Rules require Epoch AI leaderboard (Tiers 1-3 only) to show any Grok model at >=30%; primary source is EpochAI data. Current leaderboard and reports show top models (GPT-5.4 Pro etc.) at 40-52% but no Grok entry at or above 30% (Grok 4 evaluations note API issues/timeouts and low/zero scores on related tiers; May 2026 review update pending but no qualifying Grok result posted). With ~5 weeks to deadline and no credible reporting of a qualifying Grok score, No is the unambiguous resolution path.
Smart money (ranked profitable whales meowtai on Yes) leans Yes at 0.93 confidence, directly contradicting the analyst's Buy No call; these are not unranked MM proxies. Low $12.7K volume, thin book, and recent Yes price drift down (-5pp 1d) show no aggressive flow closing any alleged mispricing. No public Epoch leaderboard evidence of Grok >=30% on Tiers 1-3 as of May 2026, but whale positioning and stale post-Feb 28 trading prices indicate the market microstructure does not support the No edge.
📝 Post-mortem
Category: Mispricing
{'cause_category': 'smart_money_signal_ignored', 'explanation': "The analysis approved No despite noting 'Smart money (ranked profitable whales meowtai on Yes) leans Yes at 0.93 confidence' and 'no credible reporting of a qualifying Grok score'; meowtai held 220 shares on Yes. A Grok model reached >=30% on Epoch AI FrontierMath Tiers 1-3 leaderboard (updated June 2026), resolving Yes.", 'references': ['https://polymarket.com/event/xai-grok-score-on-frontiermath-benchmark-by-june-30', 'https://polyinsider.io/de/markets/xai-grok-score-on-frontiermath-benchmark-by-june-30'], 'key_lesson': 'Always overweight or defer to high-PnL whales on the opposite side unless you have a clear edge on their info.', 'judge_model': 'x-ai/grok-4.3'}
See today's open picks
+2 more open picks · full 3-judge reasoning · Telegram premium channel.
Subscribe Now