🔒 Hash-sealed before resolution
This prediction was committed to a SHA-256 hash at call time. Anyone can verify the call was not changed after the fact.
Will Google Gemini score at least 60% on the FrontierMath Benchmark? | No | 2026-06-30
3f7e362f
Verifying…
Verify it yourself in your terminal
echo -n "Will Google Gemini score at least 60% on the FrontierMath Benchmark? | No | 2026-06-30" | shasum -a 256 | cut -c1-8
🧑⚖️ AI judges
Gemini models remain at 36-39% on Epoch's Tiers 1-3 leaderboard as of June 2026 with no releases or jumps approaching 60% in the final 18 days; base rates and recent incremental progress make a sudden leap to the threshold implausible. The analyst's Feb 28 deadline reference is erroneous but does not change the outcome direction. Non-behavioral scientific benchmark with no gatekeeper incentives to fabricate results favors No at near-certainty.
The resolution rules require a Gemini model to hit >=60% on Epoch AI's official FrontierMath Tiers 1-3 leaderboard by June 30 2026 11:59 PM ET; current Epoch data and all credible reports show Gemini models at 36-38% with no qualifying update or release that would reach the threshold in the remaining 18 days. The analyst's cited Feb 28 trading deadline is irrelevant to the actual event window, but the factual outcome (no 60% score) still supports No. No resolver updates or evidence gaps exist that would allow a Yes resolution.
No at ~98¢ with $39k volume and tight implied spread reflects real consensus; latest Epoch leaderboard shows Gemini models at 37-39% on Tiers 1-3 with no 60% result reported as of June 2026. Smart-money whales (0.88 confidence) are concentrated on No, consistent with price action and no recent breakthroughs. Analyst deadline error does not change the outcome given the June 30 cutoff and current scores.
📝 Post-mortem
Category: Mispricing
{'cause_category': 'news_event_post_call', 'explanation': "FrontierMath v2 released June 12 2026 (exact predicted_at date) after audit corrected errors in 42% of problems, inflating Tiers 1-3 scores across models (e.g. leaders to 85%, prior ~37-48% jumps of +20-30 pts); Gemini models then qualified for >=60%. Judges' analysis relied exclusively on v1 Epoch data showing 36-39%.", 'references': ['https://www.digitalapplied.com/blog/epoch-frontiermath-v2-error-corrected-ai-benchmark-analysis', 'https://x.com/EpochAIResearch/status/2065488154086568445'], 'key_lesson': 'Always check for pending benchmark audits/revisions before assuming static leaderboards.', 'judge_model': 'x-ai/grok-4.3'}
See today's open picks
+2 more open picks · full 3-judge reasoning · Telegram premium channel.
Subscribe Now