Microsoft AI-901 (Azure AI Fundamentals)

Information Extraction Concepts

10 free practice questions with explanations

10 free questions · instant explanations · no sign-up

PassNova has 10 free Microsoft AI-901 (Azure AI Fundamentals) practice questions on Information Extraction Concepts, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.

Sample questions

Information Extraction Concepts: example questions & answers

10 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.

  1. The learning path describes information extraction as a workload that combines multiple AI techniques to extract data from content. Which two elements does it list for a comprehensive solution?

    • ASpeech recognition, and synthesis of the extracted values as audio
    • BSentiment analysis, and summarisation of the extracted document text
    • CImage classification, and semantic segmentation of document pixels
    • DText detection with OCR, and mapping the extracted text to data fields✓

    Answer: Information extraction combines computer vision, which detects text in image-based content using OCR, with machine learning or increasingly generative AI that semantically maps the extracted text to specific data fields. The expense-claim example extracts vendor, date, subtotal, tax and total from a scanned receipt in exactly this way. Speech, image classification and sentiment techniques are separate AI workloads that do not turn a document into labelled fields.

  2. What does optical character recognition (OCR) produce when it processes scanned documents, photographs of documents or PDF files that contain images of text?

    • ALabelled, structured field values
    • BSummarised, abstractive descriptions
    • CBounding boxes, but no text content
    • DEditable, searchable text data✓

    Answer: OCR automatically converts visual text in images into editable, searchable text data. Mapping that text to labelled business fields is the separate job of field extraction, which consumes the OCR output. OCR does record positional metadata such as bounding boxes, but always alongside the recognised text, and it does not generate summaries or descriptions.

  3. A batch of forms was scanned slightly rotated, so the lines of text are not horizontal. Which OCR preprocessing technique corrects this before text detection begins?

    • ALayout analysis
    • BNoise reduction
    • CSkew correction✓
    • DContrast adjustment

    Answer: Skew correction detects and corrects document rotation so that text lines are properly aligned horizontally. Techniques include the Hough transform, projection profiles and regression CNNs that predict the rotation angle directly from image features. Noise reduction removes dust spots and scanning artefacts, contrast adjustment sharpens the difference between text and background, and layout analysis belongs to the later text region detection stage.

  4. A law firm is planning an extraction solution for contracts where mistakes are unacceptable. According to the learning path's guidance on choosing an approach, what might such critical applications need?

    • AReduced model complexity for latency
    • BHuman-in-the-loop validation✓
    • COptimised system hardware
    • DCloud-based scalability

    Answer: Accuracy requirements are one of the document characteristics to weigh when planning an information extraction solution, and critical applications might need human-in-the-loop validation. Optimised hardware is the answer to high-volume processing, cloud-based solutions address scalability for variable workloads, and limiting model complexity is a response to real-time latency requirements rather than to accuracy.

  5. A developer's OCR output already contains every word on an invoice, but the accounts system still needs to know which value is the invoice number and which is the total. Which process supplies this meaning?

    • ASkew correction
    • BField extraction✓
    • CCharacter recognition
    • DText region detection

    Answer: Field extraction maps individual text values from OCR output to specific, labelled data fields that correspond to meaningful business information. OCR says what text exists in a document, while field extraction says what that text means and where it belongs in business systems. Text region detection and character recognition are stages inside the OCR pipeline that find and read the text to begin with, and skew correction is a preprocessing step that straightens rotated scans.

  6. An organisation processes a single standardised form with fixed field positions and anchor keywords, and it wants fast, explainable results without training a model. Which field detection approach fits best?

    • AMachine learning-based detection
    • BTemplate-based detection✓
    • CMulti-modal learning detection
    • DGenerative AI-based detection

    Answer: Template-based detection relies on rule-based pattern matching, using predefined layouts with known field positions and anchor keywords, label-value searches and regular expressions. Its advantages are high accuracy for known document types, fast processing and explainable results, at the cost of manual template creation and fragility when layouts vary. Machine learning, generative AI and multi-modal learning approaches are learned methods better suited to varied or complex layouts.

  7. A team supplies a large language model with the text of a document together with a schema definition and asks it to match the text to the fields in that schema. Which generative AI field detection technique is this?

    • ASupervised learning
    • BChain-of-thought
    • CPrompt-based extraction✓
    • DFew-shot learning

    Answer: Prompt-based extraction gives an LLM the document text and a schema definition so that it matches the text to the schema's fields directly. Few-shot learning instead trains or steers the model with a minimal set of examples to extract custom fields, and chain-of-thought reasoning guides the model through step-by-step field identification logic. Supervised learning is a classical machine learning approach trained on labelled datasets with known field locations, not a generative AI technique.

  8. Field extraction consumes positional metadata such as bounding box coordinates and reading order from the OCR output, not just the raw text. Why does the learning path say position matters so much?

    • ABounding boxes are needed to convert dates into a standardised ISO format
    • BCoordinates are required to correct skew before any characters can be recognised
    • CThe position of a value helps determine its meaning, such as invoice number versus phone number✓
    • DThe position of a character, not its shape, determines the OCR engine's confidence score

    Answer: Field extraction relies heavily on where text appears in a document, because the same string such as '12345' could be an invoice number, a customer ID or a phone number depending on its location. Confidence scores come from how certain the recogniser is about each character's identity, and skew correction happens during OCR preprocessing before field extraction starts. Date normalisation to ISO format works on the extracted value and does not depend on bounding boxes.

  9. An extraction solution checks that the line item subtotals on an invoice add up to the extracted invoice total before accepting the result. Which confidence and validation technique is this?

    • APattern matching confidence
    • BCross-field validation✓
    • CContext validation
    • DOCR confidence

    Answer: Cross-field validation checks the relationships between extracted fields, such as verifying that line item subtotals sum to the overall invoice total. OCR confidence is simply inherited from the underlying text recognition, and pattern matching confidence scores how well an extraction matches an expected pattern such as a date format. Context validation checks that an individual value makes sense in the document's context rather than testing arithmetic between fields.

  10. Extracted dates arrive as MM/DD/YYYY on some documents and DD-MM-YYYY on others, and the solution converts all of them to a standardised ISO format. Which stage of the field extraction pipeline performs this?

    • AField mapping and association
    • BField detection and candidate identification
    • CData normalisation and standardisation✓
    • DIntegration with business systems

    Answer: Data normalisation and standardisation transforms raw extracted values into consistent formats, such as expressing every date in the same ISO format. Date normalisation covers format detection, parsing and ambiguity resolution, alongside currency, numeric and text standardisation. Field detection finds candidate values in the OCR output and field mapping associates them with schema fields, both before the values are cleaned up, while integration is the final stage that maps the standardised fields into database schemas, API payloads or message queues.

Start practising Information Extraction Concepts →