
Coding in Agent Mode: From Idea to Shipping with GitHub
AI coding assistants have fundamentally changed how code is created. But for most developers, a chat interface has been a "read-only" or "copy-paste" coding experience — until now.

Agents are being asked to do a wide variety of work. This means a single, performance-only agentic leaderboard is no longer sufficient. Two questions decide which model to use: what’s best for my specific kind of task, and what will it cost to get that task done? We aim to help answer these questions with our latest update to Agent Arena.

At Arena, our evaluations are dynamic and grounded in real-world use. But real-world signals take time to collect. Today, we’re introducing AutoEval scores to provide immediate, calibrated model ratings on real tasks when waiting for human votes to accumulate.

Factuality remains one of the most persistent questions users face when using AI models. Today, we are launching a leaderboard that ranks models not only by human preference, but also by the factual accuracy of their responses.

Code Arena has evolved beyond frontend prototyping into a complete fullstack AI development platform. Explore capabilities like databases, authentication, third-party integrations, and deployments, and see how Fullstack Code Arena helps developers, entrepreneurs, and AI labs build, deploy, and evaluate production-ready applications.

Arena achieves $100M annualized run rate in eight months, driven by 10M+ monthly users who have contributed 700M+ conversations and 82M+ votes to build the world's largest human-preference dataset for AI evaluation.

Agents are increasingly doing real work. The resulting task distribution has greatly expanded. We desire an agent evaluation that scales along with usage and capability.

The future of AI is not single-modality chat; it is in powerful agentic capabilities. Today, we are excited to introduce Agent Mode, designed to help everyone from everyday users looking to get more done to entrepreneurs looking to maximize agentic efficacy across complex use cases.

AI coding models are increasingly used to build web apps, but aggregated leaderboards obscure key performance differences. After analyzing 250k+ Code Arena prompts, we identified major front-end task categories and built new leaderboard views to compare model strengths and weaknesses.

Max, Arena's model router powered by 5M+ community votes, is now multimodal. Starting today, Max will be available as the default option in direct chat for all modalities, with expanded capabilities including search, vision, image generation, image editing, and front-end coding. Similar to our original Max for text, the multimodal variants are latency-controlled to provide a fast and performant experience. Try it now at arena.ai/max!

For almost three years, Arena has been publishing leaderboards covering frontier AI capabilities across 10 arenas, dozens of categories, and hundreds of models, and today we're releasing the entire history of those leaderboards as a public-access dataset.

March 2026 brought major updates to the Arena leaderboard, including new rankings across document, video, text, and code models. In this monthly roundup, we break down the best AI models, latest LLM benchmarks, and key trends shaping AI evaluation.

AI failures like hallucinations are well documented. A less examined problem is that models will accept nonsensical premises without question and produce confident, detailed answers to questions that have no valid answer. BullshitBench measures whether models challenge broken premises or play along. We tested over 80 models from all major providers. Clear pushback rates range from 2% to 91%.