Cohere Parse, a vision language model built for converting enterprise documents into structured, machine-readable data, is now available in the Microsoft Foundry model . Parse joins Cohere’s Embed and Rerank models in Foundry.
What Parse does
Cohere Parse takes complex, multimodal files and returns clean Markdown suitable for downstream indexing and application logic. According to Cohere Parse can do the following
Goes beyond text recognition – Parse detects and interprets tables, forms, diagrams, and embedded images, extracting semantic context rather than a flat character stream.
Preserves document structure- The model returns bounding boxes for visual elements, which supports retrieval, grounding, and citation in downstream applications.
Is trained on enterprise document types- Cohere trained Parse on business documents across domains including finance, insurance, and scientific research and supports nine major world languages.
Is built for high-throughput workloads.- Cohere reports a measured rate of 4.5 pages per second per GPU, or approximately 36 pages per second on an 8× H100 node, served with vLLM.
For more information about this model read this blog from Cohere. https://cohere.com/blog/parse
Pricing
Cohere Parse v5 – $1.50/1K pages
Deployment and availability
Parse is available in the Foundry as a direct from Azure model. Because Parse deploys into your own Foundry environment, inference stays within your Azure tenant boundary and inherits your existing network, identity, and data governance controls relevant for teams in regulated industries with residency or isolation requirements.
Check out Cohere Parse here https://ai.azure.com/catalog/models/Cohere-parse-v5
