Microsoft Phi
Microsoft
Phi-4-reasoning-vision-15B is a compact open-weight multimodal specialist for math, science, charts, and UI grounding. Microsoft’s published results are strong for its size, but they remain provider-run and do not make Phi a broad frontier generalist.
CURRENT MODEL SNAPSHOT
Provider: Microsoft Research
Current anchor: Phi-4-reasoning-vision-15B
Lifecycle: Current
Weights: Open-weight
Reviewed: 21 July 2026

Phi’s advantage is specialization per parameter
Microsoft released Phi-4-reasoning-vision-15B on 4 March 2026. The 15-billion-parameter open-weight model combines text and image reasoning, with particular emphasis on mathematics, science, charts, diagrams, and user-interface grounding. It can switch between reasoning and non-reasoning behavior rather than forcing long traces for every input.
Phi continues Microsoft’s small-language-model strategy: concentrate training and data quality so that a compact model can perform tasks that normally require a much larger checkpoint. The practical payoff is a smaller memory footprint, lower serving cost, and more realistic private or edge-adjacent deployment.
Compact does not mean universally capable. The model’s strongest evidence is domain-shaped, and its UI understanding should not be confused with permission to execute arbitrary GUI actions without controls.
The published results are strong for 15B, but provider-run
Microsoft reports 84.8 on AI2D, 83.3 on ChartQA, 64.4 on HallusionBench, and 44.9 on MathVerse in its reruns. The team published technical detail and evaluation logs, improving inspectability, but the measurements are still produced by the model provider.
Those tasks align with the product thesis: structured visual reasoning, charts, diagrams, and mathematical content. They do not establish frontier performance on open-ended research, long-running agents, multilingual operations, or broad business judgment.
The model should be compared with other compact multimodal systems, not only 100B-plus frontier models.
UI grounding must be tested for coordinate accuracy, stale-state recovery, and destructive-action prevention.
Reasoning and non-reasoning modes should be benchmarked separately for latency and cost.
The enterprise case: a focused model for visual reasoning
Phi is compelling for chart and diagram analysis, UI understanding, scientific documents, and private multimodal workflows where a 15B model is operationally attractive. It is not the right sole model for every enterprise agent; routing specialized visual tasks to Phi can be more credible than asking it to replace a broad frontier model.
How we would evaluate it
Build a test set from real charts, scanned tables, interfaces, and diagrams, including adversarial labels and changed UI state. Measure task accuracy, hallucinated elements, coordinate errors, recovery, latency, memory, and reviewer effort. Require approval before any consequential action.
Evidence used
Beam AI support status
Under evaluation. This page does not confirm a Beam integration, supported runtime, or permission model for GUI actions.
Use case 1
Chart and diagram analyst
Test real business charts, scientific figures, and diagrams with misleading legends and missing labels. Score numerical accuracy, unsupported claims, citations to visual regions, latency, and review time.
Use case 2
UI grounding assistant
Evaluate understanding of application screenshots and forms under changed layouts, disabled controls, and stale state. Measure element identification, coordinate accuracy, recovery, and blocked destructive actions.
Use case 3
Private scientific document review
Run the 15B model on controlled technical documents with equations and images. Track extraction, reasoning accuracy, omissions, hallucinations, memory use, and reviewer corrections.
Use case 4
Compact-model router
Route visual reasoning tasks to Phi and broader research to another model. Measure routing accuracy, end-to-end completion, blended latency and cost, and whether specialization reduces review debt.

Related LLMs
Curated alternatives to compare before selecting a model family.
FAQs
Frequently Asked Questions
Model selection, deployment, governance, and Beam support questions answered.
What is Phi-4-reasoning-vision-15B?
It is Microsoft’s compact 15B open-weight model for multimodal reasoning, especially math, science, charts, diagrams, and UI grounding, with reasoning and non-reasoning behavior.
How strong are Phi’s visual benchmarks?
Microsoft reports 84.8 on AI2D, 83.3 on ChartQA, 64.4 on HallusionBench, and 44.9 on MathVerse. These are provider-run results with published technical context, not independent production evidence.
Is Microsoft Phi open-weight?
Yes. The model weights are published for download. Verify the exact license, repository, model card, notices, runtime dependencies, and artifact checksum before deployment.
Can Phi control enterprise software safely?
Visual grounding can support UI agents, but it does not make autonomous action safe. Test coordinates, stale state, permission boundaries, confirmation, destructive-action blocking, logging, and recovery.
Does Beam support Phi-4 reasoning vision in production?
Beam support is currently Under evaluation. No approved runtime, GUI-control pattern, or Beam benchmark is attached to this record.





