Jul 31, 2025

4 minutes read

# Introducing Command A Vision: Multimodal AI built for business

Command A Vision excels across enterprise image understanding tasks while keeping a low compute footprint.

Today, we're introducing Command A Vision, a new state-of-the-art generative model that brings enterprises leading performance across multimodal vision tasks while maintaining strong text capabilities. Command A Vision lets agents see inside the enterprise, unlocking the automation of tedious tasks that use visual data like slides, diagrams, PDFs, and photos. Whether it's interpreting product manuals or analyzing real-world scenes for risk detection, the model excels at tackling the most demanding enterprise vision challenges.

It surpasses other models in its class including GPT 4.1, Llama 4 Maverick, Mistral Medium 3 (and Pixtral Large) on key multimodal benchmarks. Command A Vision prioritizes enterprise needs with highly secure, efficient, and flexible deployment options. Its low serving footprint enables seamless on-premise or private deployments with two or fewer GPUs, ensuring enterprise-ready scalability.

##### **Strong performance across enterprise vision tasks**

###### **Chart, graphs, diagrams analysis**

Command A Vision excels at understanding and analyzing a wide range of visual and multilingual data, including charts, graphs, tables, and diagrams. The model accurately extracts data from diverse visual formats, applies domain-specific knowledge across industries such as finance, healthcare, manufacturing, construction, and energy, and performs complex analysis based on the extracted information.

_Methodology and further details are provided in the footnote_

###### **Document OCR and visual processing**

Command A Vision stands out in document OCR and visual processing, accurately extracting text and information from various document types, including scanned documents, invoices, and forms. The model goes beyond simple text recognition, understanding document layout and structure to extract meaningful data. This capability, combined with our structured data output support for JSON mode, allows Command A Vision to automate repetitive document processing tasks, improve data accuracy, streamline workflows, and seamlessly integrate with existing systems, making it an invaluable tool for enterprises processing large volumes of documents. It achieves top-tier performance across the DocVQA, TextVQA, and OCRBench benchmarks.

_Methodology and further details are provided in the footnote_

###### **Real-world scene understanding**

Command A Vision's capabilities extend to real-world scene understanding, enabling it to analyze and interpret complex visual environments. This goes beyond simple object detection, as the model can understand spatial relationships, context, and even subtle nuances within images and photographs, making it ideal for a wide range of real-world applications, including risk detection in industrial settings and retail analytics.

##### **Capabilities and efficiency suited for enterprise scale**

Command A Vision was built to serve enterprises across the capabilities that matter most to them. It combines the other important text features of Command A, like advanced retrieval-augmented generation (RAG) with citations and multilingual performance across several key business languages.

With Command A Vision, enterprises can quickly and securely access context-aware insights and analysis on their own data, whether in text or various types of enterprise image formats – on the hardware they have and in the languages they need. Our Command model series is optimized for enterprise needs at the forefront, excelling on complex business applications while balancing performance, accuracy, and efficiency.

With low hardware requirements, enterprises in regulated industries that need private deployments can efficiently use Command A Vision in production. Command A Vision can be deployed privately with just two or fewer GPUs. It only requires two A100s, or one H100 for 4-bit quantization.

##### **What customers are saying**

> _“We’re incredibly excited about the release of Command A Vision. These models dramatically expand the boundaries of what’s possible with generative AI, enabling us to move beyond text and into the realm of visual understanding. Already, we’ve seen Command A Vision solve some of our most complex and time-consuming challenges; it not only streamlines workflows but unlocks entirely new opportunities for generative AI. By integrating visual context into our AI systems, we can start building solutions that are grounded in what we can see, not just what we can read. I’m excited to see how far we can push this technology and what we can accomplish with it in our toolkit.” – Jeffrey English, Director, Professional Services, Fujitsu Intelligence_

> _“During early testing, the Command A Vision model has demonstrated exceptional capabilities in understanding and extracting data from intricate construction industry documents, such as lien waivers, invoices, and drawings. The ability to automate this kind of AI-driven data capture has the power to transform document processing, data accuracy, and project management that could reduce risk, time, and cost for the construction industry.” – Mark Webster, Senior Vice president and General Manager, Oracle Infrastructure Industries_

##### **Availability**

Command A Vision is available today on the [Cohere platform](https://dashboard.cohere.com/welcome/login?redirect_uri=%2Fplayground%2Fchat%3Fmodel%3Dcommand-a-vision-07-2025) and for research use on [Hugging Face](https://huggingface.co/CohereLabs/command-a-vision-07-2025?ref=cohere.com/blog). If you are interested in private or on-prem deployments, please contact our [sales team](/content/contact-sales/index.html) for bespoke pricing.

\[1\] We compare ourselves to the best non-reasoning models (non-deprecated) available from other providers. Where available, benchmark scores are taken either from other providers’ own reports or publicly available leaderboards; or otherwise (greyed numbers), use best-effort internal evaluation either through VLMEvalKit or the official codebase.
