Gemini 3.6 Pro
Google DeepMind
Frontier Complex Reasoning & Multi-Step Mathematical Proofs
Achieves SOTA on GPQA Diamond (89.4%) and MATH-500 (94.2%) with verified multi-step reasoning trace verification.
Every major AI capability milestone, benchmark result, and model release — tracked, scored by capability index, and linked to original sources. Updated with each verified release.
Google DeepMind
Achieves SOTA on GPQA Diamond (89.4%) and MATH-500 (94.2%) with verified multi-step reasoning trace verification.
Anthropic
Independent leaderboard score of 69.2% on SWE-bench Pro under standardized harness conditions without scaffold exploits.
DeepSeek AI
Matching proprietary frontier reasoning models on MATH and Codeforces while releasing full open weights and training logs.
GLEE Core OS
Autonomous AI Operating System coordinating heterogeneous LLMs (Zoro, Sanji, Robin, Sabo) with zero single-point hallucination.
OpenAI
Demonstrates 99.8% accuracy on Needle-in-a-Haystack across 1M+ tokens with full decontaminated evaluation.
Meta AI
First 400B parameter open multimodal model natively processing high-FPS video streams and 3D spatial representations.
GLEE Core Platform
Distributed multi-agent orchestration engine delivering zero-conflict task routing and consensus protocol verification across heterogeneous AI foremen.
Google DeepMind
Ultra-low latency end-to-end multimodal audio processing allowing natural conversational interruption, emotional inflection, and live video analysis.
GET /agora/api/capabilities
?category=reasoning&sort=score
Full JSON export of all capability entries. Filter by category, sort by score or date. Used by GLEE's own AI agents to stay current on the frontier.