← Ask a question

What are the latest developments in artificial intelligence research and frontier models?

Full research answer

The frontier AI landscape is moving at an extraordinary pace heading into mid-2025, with the major labs shipping significant model updates, new reasoning architectures going mainstream, and the competitive gap between open and closed models narrowing faster than most analysts predicted a year ago.

Google DeepMind made a substantial splash with Gemini 2.5 Pro, which currently sits atop several major benchmarks including MMMU and GPQA Diamond, outperforming OpenAI's o3 on a number of reasoning tasks. Notably, Gemini 2.5 Pro achieved a 18.8% score on Humanity's Last Exam — a deliberately brutal benchmark designed to stump frontier models — marking one of the highest recorded scores on that test. Meanwhile, OpenAI pushed o3 and o4-mini into broader availability, with o4-mini surprising many researchers by punching well above its weight class on math and coding tasks relative to its inference cost. OpenAI also previewed GPT-4.1, positioned as a more efficient coding-focused model with a 1 million token context window now available via API.

On the open-weight side, Meta's Llama 4 family launched with the Scout and Maverick variants, with Maverick claiming competitive performance against GPT-4o on MT-Bench while remaining fully open. This release intensified the ongoing debate about whether open models can meaningfully keep pace with closed frontier systems — and for many enterprise use cases, the answer is increasingly yes. Separately, Mistral released Mistral Small 3.1, a 24B parameter model with strong multilingual performance and a 128k context window, further populating the mid-tier open model space. On the research side, a wave of papers around "test-time compute scaling" — the idea of spending more inference-time compute to improve output quality — has become arguably the dominant architectural conversation, building on the foundations that OpenAI's o-series and Google's Gemini thinking models established.

Watch for Anthropic's next move most closely. Claude 3.7 Sonnet's extended thinking mode has been widely praised for coding and agentic workflows, but the lab has been quieter than its peers on new releases, suggesting something significant may be staged. The broader theme to track is the shift from raw benchmark performance to real-world agentic capability — how well these models execute multi-step tasks autonomously. That's where the next competitive frontier is being drawn, and the gap between "impressive demo" and "reliable deployment" remains the defining challenge of 2025.

— Arcade

Get this kind of insight every morning

Arcade and the Lumis research team brief you on everything that matters — before you start work.

Subscribe free →