Concept and mechanism
Image classification answers about overall content; object detection adds instance locations. OCR extracts characters from images. Layout analysis retains structure such as tables, rows, and positions, useful when an amount must remain associated with the correct entity. Document Intelligence offers reading, layout, and document-extraction models. Face detection identifies presence and position without proving who the person is. Identity recognition is a distinct capability with its own requirements and availability. Choose the output the process needs before choosing a service name.
Guided application
In a fictional financial integration, include low-resolution documents, split tables, and different decimal separators. Compare extracted values with the source and control totals. Confidence helps route cases but does not guarantee an individual result is correct or authorize posting. Define review for discrepancies and measure errors by document type. Handover should identify model, version, supported layout, and exception handling. There is also a lifecycle decision: Image Analysis 4.0 is deprecated with retirement announced for 2028-09-25. That does not mean every Vision service is retiring. For new projects, check alternatives and support throughout the maintenance horizon.
A high-confidence extracted value remains pending if it contradicts the document currency or total.
Common pitfalls
OCR as a structured table; detected face as identity; confidence as proof; historical syllabus as current support.
Related topics: Workloads and operational responsibility · Machine learning and useful evaluation · Language, speech, and operational meaning
Visual output needs business validation and a supported dependency.
Reference: Document Intelligence extraction and layout · AI-900 historical skills measured 2025-05-02; exam retired 2026-06-30