Microsoft AI-901 (Azure AI Fundamentals)

Vision & Image Generation with Foundry

14 free practice questions with explanations

14 free questions · instant explanations · no sign-up

PassNova has 14 free Microsoft AI-901 (Azure AI Fundamentals) practice questions on Vision & Image Generation with Foundry, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.

Sample questions

Vision & Image Generation with Foundry: example questions & answers

14 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.

  1. Which term does the unit use for models that combine visual understanding with natural language responses, such as GPT-4.1 in Foundry?

    • APurpose-built analyzers
    • BText-to-image generation models
    • CVision-enabled GPT models✓
    • DText-to-video models

    Answer: The ability of a model to combine visual understanding with natural language responses is referred to as vision-enabled GPT models or GPT with vision, and such models are designed for flexible, general-purpose visual reasoning. Text-to-image and text-to-video models generate visual content rather than interpreting it, and purpose-built analyzers belong to Azure Language for deterministic text analysis.

  2. A company is building a production-grade AI agent that must reason over images and text, return structured outputs and use tools. Which model family in the Foundry catalog does the unit describe as designed for these enterprise and agentic scenarios?

    • AThe GPT-5 series, such as GPT-5.1 and GPT-5.2✓
    • BThe GPT-Image-1 family, such as GPT-Image-1.5
    • CThe Sora family, such as Sora 1 and Sora 2
    • DThe GPT-4.1 family, such as GPT-4.1-nano

    Answer: The GPT-5 family in Foundry, including GPT-5.1 and GPT-5.2, comprises advanced multimodal models that support text and image inputs, structured outputs, tool use and large-context reasoning, and they are typically used in production-grade agents. GPT-4.1 and its mini and nano variants are general-purpose multimodal models used for description and analysis tasks, GPT-Image-1 models generate images, and Sora models generate video.

  3. According to the unit, in which two forms can an image be supplied alongside a text prompt in the Responses API?

    • AAs a storage container name or as an account key
    • BAs a URL or as base64-encoded image data✓
    • CAs a PNG attachment or as an SVG string
    • DAs a byte stream or as a local file path

    Answer: With the OpenAI Responses API in Foundry, a single request can include text and image input together, and images can be provided as URLs or as base64-encoded image data. Base64 converts binary image bytes into safe ASCII text so they can be embedded in JSON. The unit does not describe supplying images as container names, account keys, SVG strings, byte streams or local file paths.

  4. In the image-analysis code sample, what is the value of the 'type' field for the content item that carries the picture?

    • Ainput_image✓
    • Binput_text
    • Cinput_audio
    • Dimage_url

    Answer: The user message content in the sample is a list containing an item with "type": "input_text" for the question and an item with "type": "input_image" whose image_url field holds the picture. image_url is the name of the field inside that item rather than its type, input_text is the type of the text instruction, and input_audio is not used in the sample.

  5. Which image generation model does the unit say Microsoft especially recommends starting with for most new projects?

    • AGPT-Image-1-Mini
    • BGPT-Image-1.5✓
    • CGPT-4.1-mini
    • DGPT-Image-1

    Answer: For most new projects Microsoft recommends starting with the GPT-Image-1 family, especially GPT-Image-1.5, because of its improved quality, editing support and enterprise readiness. GPT-Image-1 is the earlier general-purpose model, GPT-Image-1-Mini is the lighter, cost-efficient variant for experimentation and high-volume work, and GPT-4.1-mini is a vision-enabled model for analysing images rather than generating them.

  6. In the image generation code sample, which argument tells the Responses API to produce an image from the prompt?

    • Atools=[{"type": "input_image"}]
    • Btools=[{"type": "image_generation_call"}]
    • Ctools=[{"type": "image_generation"}]✓
    • Dresponse_format="b64_json"

    Answer: The sample calls client.responses.create with tools=[{"type": "image_generation"}], which enables the image generation tool for the deployed gpt-image model, and it then reads the base64 result from the output item whose type is image_generation_call. image_generation_call is the type of the returned item rather than the tool, input_image is the content type for supplying a picture as input, and response_format is not an argument the sample passes.

  7. How does the image generation sample obtain the picture data from the response before writing foundry_generated.png?

    • AIt polls a video id endpoint until the status is completed and fetches the content
    • BIt downloads the file from the URL in response.output_text and saves the bytes
    • CIt reads response.data[0].b64_json and writes the string straight to disk
    • DIt base64-decodes the result of the output item whose type is image_generation_call✓

    Answer: The sample uses next(item.result for item in response.output if item.type == "image_generation_call") to get the base64 string, then writes base64.b64decode(image_base64) to foundry_generated.png. The response does not return a download URL in output_text, the sample does not read a data list with b64_json, and polling a video id endpoint is the Sora video generation pattern.

  8. Which third-party family of open-source image generation models from Black Forest Labs does the unit say you can also access in Foundry?

    • AFLUX✓
    • BLlama
    • CPhi
    • DSora

    Answer: FLUX is a family of open-source image generation models created by Black Forest Labs, designed to produce high-quality, photorealistic and stylistically flexible images from text prompts, and it is available as a third-party model in Foundry. Sora is OpenAI's video generation model, and neither Llama nor Phi is the Black Forest Labs image family the unit names.

  9. A marketing team wants to remix an existing video with targeted edits and have audio generated, rather than regenerating the whole clip. Which Foundry model supports this?

    • ASora 2✓
    • BGPT-5.2
    • CSora 1
    • DFLUX

    Answer: Sora 2, in public preview, supports text-to-video, image-to-video and video-to-video (remix), and it introduces audio generation, improved realism and remixing capabilities that allow targeted edits instead of regenerating an entire video. Sora 1 generates short clips from text or images without these remix and audio features, while GPT-5.2 and FLUX do not generate video output.

  10. Which statement about video generation models in Foundry is correct?

    • AGPT-5 series models can generate video output when they are deployed with the video tool enabled
    • BFLUX models generate both images and short video clips from the same text prompt
    • CGPT-Image-1.5 can produce video when the seconds parameter is included in the request
    • DSora models are the only native video generation models provided directly through Foundry✓

    Answer: Sora models are currently the only native video generation models provided directly through Foundry; other Foundry models may be multimodal across text, image and audio, but they do not generate video output. GPT-5 series models reason over text and images, FLUX and GPT-Image-1.5 generate still images, and no video tool or seconds parameter turns them into video generators.

  11. A developer's script has just POSTed a video generation request to the /openai/v1/videos endpoint with model sora-2. What should the script do next?

    • ASend the same request again with a larger seconds value
    • BCall the /videos/{video_id}/content endpoint at once
    • CPoll /videos/{video_id} until the job status is completed✓
    • DRead the MP4 bytes from the body of the POST response

    Answer: Video generation is resource-intensive and runs as an asynchronous job: you create the job, poll for its status, and download the video only once the job is complete, which often takes one to five minutes. The POST response returns a video id to poll rather than the MP4, the content endpoint with variant=video is called only after the status is completed, and resubmitting the request would simply start another job.

  12. In the Sora 2 REST example, which two request body parameters control the resolution and length of the generated clip?

    • Aquality and duration
    • Bsize and seconds✓
    • Cwidth and frames
    • Dresolution and length

    Answer: The curl body in the sample sets "model": "sora-2", a "prompt", "size": "1280x720" for the resolution and "seconds": "8" for the clip length. The playground describes these as video dimensions and duration, but the parameter names in the request are size and seconds rather than resolution, width, quality, length, frames or duration.

  13. In which Foundry portal feature can you experiment with Sora models by describing a video and specifying parameters like dimensions and duration?

    • AThe Foundry Video Playground✓
    • BThe Language Playground
    • CThe Content Understanding page
    • DThe Voice Live playground

    Answer: Sora 1 and Sora 2 are exposed through the Azure OpenAI API and the Foundry Video Playground, where you deploy a video generation model, describe the content you want and set parameters such as video dimensions and duration. The Language Playground is for Azure Language features, the Voice Live playground is for spoken agents, and Content Understanding is for extracting information from content.

  14. Which header does each curl call in the Sora 2 example use to pass the credential to the video endpoints?

    • A-H "api-key: $AZURE_OPENAI_API_KEY"✓
    • B-H "Subscription-Key: $AZURE_OPENAI_API_KEY"
    • C-H "Authorization: Bearer $AZURE_OPENAI_API_KEY"
    • D-H "x-api-key: $AZURE_OPENAI_API_KEY"

    Answer: Each curl call in the Sora 2 example authenticates with the header -H "api-key: $AZURE_OPENAI_API_KEY", which passes the Foundry resource key. A bearer Authorization header, a Subscription-Key header and an x-api-key header are not the header the sample sends to the /openai/v1/videos endpoints.

Start practising Vision & Image Generation with Foundry →