Mistral AI shipped OCR 4, a document intelligence tool that extracts structured content from documents with bounding boxes and confidence scores across 170 languages.
The system deploys in a single container and integrates into enterprise search and structured data pipelines. Mistral claims OCR 4 delivers 4x speed advantages over competing systems while maintaining high accuracy, particularly with low-resource languages.
Meanwhile, Anthropic introduced Claude Tag, a Slack-based workflow system that lets teams assign tasks to Claude while connecting it to tools and codebases. The AI assistant retains context across channels and has become integral to Anthropic's internal operations.
The company said its product team uses Claude Tag to generate code and assist with analytics, support, and debugging tasks. The system represents Anthropic's push into enterprise workflow automation.
ByteDance unveiled Seedance 2.5, an AI video generation model that creates 30-second, 4K videos from single prompts. Users can provide up to 50 images, videos, or audio clips as reference materials for greater control over output.
The model launches in China next month, though ByteDance has not announced international availability. The release intensifies competition in AI video generation as companies race to extend clip lengths and improve quality.
Enterprise AI Infrastructure Gains Momentum
NVIDIA's Agent Toolkit is enabling businesses to build specialized AI agents using open models and secure runtime environments. Companies including Cadence, Synopsys, and CrowdStrike are deploying the technology across life sciences, healthcare, cybersecurity, and industrial operations.
The toolkit integrates with existing enterprise tools and data systems, allowing organizations to create domain-specific agents rather than relying on general-purpose models. This shift toward specialized AI reflects growing enterprise demand for controllable, auditable AI systems.
Separately, researchers released technical insights on prompt injection vulnerabilities, highlighting how large language models struggle to distinguish between their own reasoning and external inputs. The analysis suggests current security measures remain insufficient as models process all information through unified token streams.
💬 Discussion
Sign in to join the discussion.
Sign in →No comments yet — be the first.