13 min read

The AI Model Landscape for Building Agents

By submitting, you consent to our use of your data. Privacy Policy.

Category

AI Agents

Share the article

Key takeaways

  • An agent isn't one job, so "which LLM is best" is the wrong question. Pick a model for each step.

  • Beyond LLMs, decision models, predictive models, forecasting models and optimization solvers answer with a choice, a number or a plan, not text.

  • The cost gap at the top is large: six points on Artificial Analysis's index cost about eight times as much per task.

  • Fine-tuning a small open model pays off against frontier prices at volume, not against the cheap tier.

  • Much of the specialist layer is being bought: OpenRouter by Stripe, Kumo by NVIDIA, Prior Labs by SAP, World Labs by AMD.

Every time a new model comes out, the question we ask internally is the same: where does it fit in a Beam agent? The answer is rarely "everywhere", because an agent isn't one job. It reads documents, makes decisions, takes actions and checks its own work, and in 2026 each of those jobs has models built for it.

So instead of ranking LLMs, we mapped every kind of model you can build an agent with today: seventeen categories, the companies in each, and how we think about them at Beam.

Market map of the models available for building AI agents, October 2026, in seventeen categories grouped by the job they do: reason (frontier, open-weight, small and fast LLMs), decide (decision models, predictive and tabular models, forecasting, optimization solvers), perceive (retrieval, document parsing, vision, voice), create (image generation, 3D and world models), act and check (computer use, guardrails), and run and improve (routers, distillation and fine-tuning)

The agent model landscape, October 2026: seventeen categories grouped by the job each model does inside an agent. Open the image for full size.

We placed each model once, in the job it's mainly used for, and included only models you can call or download today, plus the routers and training platforms around them. Agent frameworks and apps are out. Prices and scores are as of 5 October 2026.

Reason: the LLMs

Frontier models

The models at the top of the leaderboards. Anthropic holds the top five spots on Artificial Analysis's Intelligence Index (v4.3), with OpenAI and Google tied just behind, and it leads on spend too: 40% of enterprise LLM API spend, against 27% for OpenAI and 21% for Google (Menlo Ventures, December 2025).

Capability against cost on Artificial Analysis's Intelligence Index v4.3: Claude Opus 5.5 at max effort scores 58 at $5.98 per task, GPT-6 Astra 53 at $3.26, GPT-6.1 Sol 52 at $0.72, Claude Opus 5.5 at default effort 51 at $1.34, and Muse Spark 1.3 48 at $1.60

Capability against cost per task on Artificial Analysis's Intelligence Index, October 2026.

How we see it at Beam

They aren't interchangeable, whatever the benchmark says. In my own use, Claude is noticeably better at design work than OpenAI's models. With GPT-6.1 Sol, I had to spell out basic layout instructions I'd expect a model at that level to handle. Every model has jobs where it's stronger or weaker than its benchmark score suggests.

Open-weight models

Models you can download, host and fine-tune yourself. The three strongest are all Chinese, led by Xiaomi's MiMo-V2.6-Pro, and they sit about 12 points behind the closed frontier on the same index. They lead on usage instead: Chinese models carried 57–67% of tokens on OpenRouter in mid-September (CNBC). Developers routing for price and enterprises buying for production are making different choices.

Open-weight models take 11% of enterprise LLM API spend (Menlo Ventures, December 2025) but about 60% of tokens routed through OpenRouter from the US (The New Stack, August 2026)

Two measures, two buyers: enterprise spend and routed developer tokens.

Small, fast models

The cheap tier for light steps like classification and extraction. OpenAI's GPT-6 Luna costs $0.10 per million input tokens, and models like Gemma 4 and Liquid AI's LFM2.5 run on a laptop.

How we see it at Beam

This is the tier worth watching. The models worth moving to are usually the cheaper tier that has caught up with last year's top model, because that can cut a bill substantially without losing accuracy. We go through that decision in our piece on whether a cheaper LLM is worth switching to.

Decide: models that answer with a choice, a number or a plan

Decision models

The newest category on the map, and three weeks old. TypeSafe AI launched Jev on 15 September with a $40M seed, and within 16 days Cloudflare, OpenAI, Perplexity, AWS and Fastino had launched their own. Cloudflare describes a decision model as one that "makes classifications to help agents decide how to act, based on certain probabilities" (Cloudflare).

For high-volume steps where the answer is one of a known set, they're faster and more predictable than a general model. Every benchmark in the category so far is run by a vendor.

A general LLM answers a refund request with free text that has to be parsed; a decision model returns approve, escalate or reject with a probability, and the agent acts alone only above a threshold

A decision model returns one of a known set of options with a probability, so the agent can act alone above a threshold and ask a person below it.

Predictive and tabular models

Statistical models, which surprisingly few people talk about. You train them on a large history of past decisions, and they learn to repeat those decisions. They can be cheap and very accurate on familiar cases, and fail badly on anything new. XGBoost is still the default at about 30 million downloads a month, but 2026 brought foundation models for tables: NVIDIA bought Kumo, SAP bought Prior Labs, and Fundamental raised $255M. None of the VC maps we read put these models in the agent stack.

How we see it at Beam

We build almost everything on LLMs today, for two reasons. The first is trust: anyone can read the prompt and see what the agent has been told to do, while a statistical model is a black box. The second is coverage: a statistical model only makes sense for a decision that repeats at high volume with a long, stable history, and it doesn't make sense to build one that only two or three customers out of twenty would use.

Statistical models are cheap and very accurate on familiar cases but fail badly on new ones and are a black box; LLMs cost more per call but handle variety and let you read the prompt

Where statistical models and LLMs each win, and why Beam builds almost everything on LLMs today.

Forecasting models

Pretrained models that forecast a numeric series without being trained on it first, like Google's TimesFM-3, Amazon's Chronos-2, Datadog's Toto 2.0 and Nixtla's TimeGPT. LLMs are poor at extrapolating numbers and give no uncertainty range, so an agent planning stock or staffing calls one of these instead. On GIFT-Eval, the main forecasting benchmark, the top entries in September were systems that use an LLM to choose between forecasting models rather than any single model (GIFT-Eval).

Optimization solvers

Solvers like Gurobi, NVIDIA's cuOpt and Google's OR-Tools find the best plan under constraints, such as a delivery route, a shift schedule or an energy dispatch, and prove that it's feasible. An LLM can't search millions of combinations or certify an answer, and even describing the problem is hard: on MIPLIB-NL, built from real industrial problems, the LLMs tested wrote a correct model only 24–39% of the time (MIPLIB-NL). OSM, which optimizes energy systems, splits the work: the LLM writes the decision as a query, and a solver answers it (OSM, VLDB 2026). Gurobi and Nextmv, now part of FICO, shipped agent interfaces for their solvers this year.

An LLM asked for a plan directly returns one with no proof it is feasible or optimal; an LLM that writes the problem as a model hands it to a solver, which returns an optimal plan with feasibility evidence

The LLM describes the problem, and the solver finds the plan and proves it meets the rules.

Perceive: search, documents, images and voice

Retrieval

Embedding models and rerankers decide which few thousand tokens the agent actually sees. The independents are being bought: Voyage AI now belongs to MongoDB and Jina AI to Elastic. Cohere released Embed 5 on 30 September, and Google's gemini-embedding-2 puts text, images, video and audio in one search space.

How we see it at Beam

When an SOP fits in the model's context window, which with today's models can be up to a million tokens, we put it in the system prompt instead of retrieving it. Retrieval and file search are good but still a bit lossy: by my estimate, around 0.5% of the time the model doesn't look up the information it needed. When you're aiming for around 90% accuracy on the most complex cases, that miss rate matters, and a document in the system prompt is guaranteed to be read.

Document parsing

Parsers turn PDFs, scans and spreadsheets into structured text with tables and layout intact, at a fraction of the cost of sending every page through a frontier model. Mistral OCR 4.1 costs $4 per 1,000 pages and Reducto's new r-1 model costs 1¢ a page.

Vision

Specialist vision models return exact boxes, masks, counts and click coordinates, where a general LLM can describe an image but is weak at locating things in it. Meta's SAM 3.1, Moondream 3.1 and Roboflow's RF-DETR are the main options, and Microsoft's OmniParser turns screenshots into elements a computer-use agent can click.

Voice

The best-funded specialist category on the map. ElevenLabs was valued at $22B on 30 September and Deepgram passed $100M in annual revenue. The frontier labs now ship native speech-to-speech models too, with OpenAI's GPT-Live-1 and Google's Gemini 3.8 Live both launching in September. That leaves a real architecture choice: a cascade of speech-to-text, an LLM and text-to-speech gives you a transcript to check at every step, while a single speech-to-speech model is faster and handles interruptions more naturally.

A cascade runs speech to text, an LLM and text to speech, with a transcript at every step; native speech-to-speech uses one model and is faster with more natural turn-taking

The two architectures for a voice agent.

Create: 2D and 3D

Image generation

OpenAI's GPT Image 2.5 holds the top two spots on Artificial Analysis's text-to-image leaderboard, and Meta's Muse Image is the cheapest model in the top ten at about $10 per 1,000 images (Artificial Analysis). For marketing and design agents these are the output step: according to Google, WPP built Nano Banana into its agentic marketing platform.

3D and world models

Two groups. 3D generators like Meshy, Tripo and Tencent's HY 3D turn a prompt or a photo into a textured mesh, and the money is moving fast: Meshy passed $100M in annual revenue in September. World models generate an environment you can move through, for simulation and robotics. AMD agreed to buy World Labs for about $8.2B in September (AMD), and World Labs' API and NVIDIA's open Cosmos 3 are the ones you can use today without a sales call.

Act and check

Computer use

Models that read a screen and click, for systems without an API. This category is being absorbed into the frontier models: Google shut down its standalone computer-use model in July and built the capability into Gemini Flash. The specialists left, like H Company's Holo4 and Microsoft's Fara1.5, compete on cost and size. Vendors report different versions of the OSWorld benchmark, so their headline scores can't be compared.

Guardrails

Small classifiers that screen inputs, outputs and tool calls, like Meta's Llama Guard 4 and IBM's Granite Guardian 4.1. The independents are being bought: Check Point bought Lakera, Palo Alto Networks bought Protect AI, and Harvey bought Guardrails AI in September.

How we see it at Beam

Our checks are specific to each step. When a check fails, we work out whether the output was wrong or the input data was wrong, then either fetch the right data or send the output back with the specific mistake named.

When a step's check fails, work out whether the output or the input data was wrong: fetch the right data, or send the output back with the specific mistake named, then rerun the step

How a failed check is handled in a Beam agent.

Run and cut costs

Routers

A router picks a model for each request. Stripe agreed to buy OpenRouter in August for a reported $7.5B, and Microsoft, NVIDIA and Cloudflare now ship routers of their own.

How we see it at Beam

There's no hidden router in Beam. The model is a setting on each step of an agent: each step starts with a default from the workspace's preferred models, can have fallback models that take over if a call fails, and can retry when accuracy drops below a threshold you set. Changing a step's model is a deliberate decision, and it changes that step only. OpenAI rolled back the automatic router in ChatGPT for free users in December 2025.

A hidden router sends the same step to a different model on each run; in Beam the model is a setting on the step, with a fallback model if a call fails and a retry when accuracy drops below your threshold

A hidden router, compared with a model chosen for each step.

Fine-tuning and distillation

Fine-tuning trains an open model further on your own examples, so a small model does one narrow job as well as a frontier model. Three approaches now usually run as one pipeline: supervised fine-tuning on examples, mostly with LoRA, which trains a small adapter instead of the whole model; distillation, where a frontier model's accepted outputs become those examples; and reinforcement fine-tuning, where a grader scores the model's answers and training pushes it towards higher scores. Fireworks, Together, Baseten and Thinking Machines' Tinker run the training, and Unsloth is the open-source option.

Fine-tuning as one pipeline: a frontier model's accepted outputs become training examples, a small LoRA adapter is trained on an open base model, a grader scores answers in reinforcement fine-tuning, and the result runs on a dedicated GPU

Distillation, supervised fine-tuning and reinforcement fine-tuning, run as one pipeline.

The savings can be large. Genspark and Fireworks trained the open MiniMax M3 model this way into a slide-making model that matched Claude Opus 5 on Genspark's own evaluation, at about 90% less per finished deck (Fireworks, vendor claim). EliseAI cut inference cost by about 60%, and latency from 2.2 seconds to 250 milliseconds, with a fine-tuned 4B model on Baseten (Baseten).

There are two catches. Fireworks and Together no longer serve your own fine-tune per token, so running one means renting a GPU, about $5.50 to $8 an hour for an H100. That pays off against frontier prices at volume, not against the cheap tier. And a fine-tune of a closed model ends when its base model retires: OpenAI is winding down self-serve fine-tuning, and Claude 3 Haiku, the only Claude model ever fine-tunable on AWS Bedrock, reached end of life in September.

How we see it at Beam

We've already started. A large frontier model produces high-quality outputs, and we use those as ground truth to bring a much smaller model up to the same standard on that task. It's the same idea as in our piece on why self-learning agents stall on weak feedback: a small model can do a lot more when a strong one shows it what good looks like.

How we use the map at Beam

Different steps in the same agent use different models on purpose: a fast, cheap model on light steps, a frontier model on the steps that need real reasoning.

One model everywhere sends every agent step to a frontier LLM; a model per step uses a frontier LLM for planning and odd cases, a small LLM for classification and extraction, a decision model to approve or escalate, and code for calculations

Assigning a model to each step of an agent.

The cost gap is the reason this matters. On Artificial Analysis's Intelligence Index, Claude Opus 5.5 scores 58 and GPT-6.1 Sol scores 52, but Opus costs about $5.98 per benchmark task against $0.72 for Sol (Artificial Analysis, October 2026). On some steps those six points are worth eight times the cost. On most, they aren't.

Once agents are easy to set up and work reliably, the next problem is cost at scale. Our plan is to start every workflow on LLMs, because they handle variety and they're easy to trust. Once a workflow has run long enough, its traces become a training set, and the repetitive parts can be distilled into something smaller and cheaper: a small model, a fixed workflow, or a statistical model where the history supports it.

Beam's path to lower cost: start every workflow on LLMs, collect the agent's traces, distil the repetitive parts into a small model, a fixed workflow or a statistical model, and keep the LLM for cases that vary; for 1 million agent calls a month, Claude Opus 5.5 costs about $18,000, GPT-6.1 Sol $9,000, a fine-tuned 9B model on one H100 about $3,950 and GPT-6 Luna $450

Fine-tuning pays off against frontier prices at volume, not against the cheap tier.

So which model is best for an AI agent? The one that's best for the step in front of you, at a cost that still makes sense when the agent runs a million times.

Common questions about models for AI agents

What kinds of models can you build an AI agent with?

Seventeen kinds, by the job they do: frontier, open-weight and small LLMs for reasoning; decision, predictive, forecasting and optimization models for answers that are a choice, a number or a plan; retrieval, document-parsing, vision and voice models for perception; image, 3D and world models for creating; computer-use and guardrail models for acting and checking; and routers and fine-tuning platforms to cut costs.

Which LLM is best for AI agents?

It depends on the step. A frontier model suits planning and judgement on unfamiliar cases, a smaller and faster model suits classification and extraction, and decision models or code suit fixed choices and calculations. Test each step on your own cases, since public rankings change monthly.

Does Beam automatically pick the best model for each task?

No. In Beam, the model is a setting on each step of an agent, with sensible defaults, fallback models if a call fails and retries when accuracy drops below a threshold. Changing a step's model is a deliberate choice, made per step.

When should you use a statistical model instead of an LLM?

When the same decision repeats at very high volume and there is a large, stable history of past decisions to train on. Statistical models can be cheap and accurate on familiar cases but fail on new ones, so LLMs remain the better choice where cases vary.

Start Today

Start building AI agents to automate processes

Join our platform and start building AI agents for various types of automations.

Start Today

Start building AI agents to automate processes

Join our platform and start building AI agents for various types of automations.