Tables, forms, diagrams and images in enterprise documents can be converted into structured Markdown for AI retrieval by Cohere's new Parse model. Parse is a vision language model designed to process large volumes of documents for indexing, retrieval-augmented generation (RAG) and agentic retrieval. It works across nine major languages.
Parse goes beyond text recognition. It detects visual elements such as tables and images and returns bounding boxes, preserving document structure for retrieval, grounding and automation.
It has been trained on business documents from industries including finance, insurance and science.
Parse is available through the Cohere API at US$1.50 per 1000 pages. It can also run in Cohere's Model Vault for single-tenant inference at a lower cost per page.
Organisations in regulated industries can deploy it on their own infrastructure, in a private cloud or on-premises, with a small serving footprint.
It is also part of Cohere's Compass search and retrieval stack, alongside its Embed and Rerank models.
Cohere says Parse scores 79.2 on ParseBench, a benchmark of parsing performance for agent use, averaged across three dimensions.
That compares with 74.5 for Mistral OCR 4, 72.4 for Databricks AI Parse and 78.3 for LlamaParse's Cost Effective offering, in Cohere's evaluation.
Cohere says Parse leads AWS Textract and Google Document AI by more than 20 points. Only much larger general-purpose frontier models scored higher.
Parse scored 87.0 on tables and 86.6 on content faithfulness, but 64.0 on semantic formatting.
Document parsing has become a crowded field as vendors target the ingestion step feeding enterprise AI. Recent entrants covered by IDM include ABBYY's FineParser and Elastic's jina-ocr-v1.