Batch & Stream Processing
12 free practice questions with explanations
12 free questions · instant explanations · no sign-up
PassNova has 12 free Microsoft DP-900 (Azure Data Fundamentals) practice questions on Batch & Stream Processing, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.
Batch & Stream Processing: example questions & answers
12 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.
What is the defining characteristic of batch processing?
- AData is queried directly from a semantic model without any transformation
- BEach new piece of data is processed individually the instant it arrives in the stream
- CMultiple data records are collected and processed together as a single batch✓
- DData is replicated continuously from a source database into a separate data lake for storage
Answer: Batch processing collects and stores records, then processes the whole group together, often on a schedule or trigger. Processing each new record immediately as it arrives describes stream processing instead. Querying a semantic model and continuous replication describe reporting and mirroring scenarios, not a batch processing style at all.
A credit card company sends customers one monthly bill that combines all of that month's purchases. Which data processing approach does this illustrate?
- AReal-time processing
- BStream processing
- CBatch processing✓
- DStructured streaming
Answer: Collecting a month's worth of transactions and processing them together into a single bill is a classic real-world example of batch processing. Stream processing and real-time processing instead process each transaction individually as it occurs, which would produce a bill per purchase rather than per month. Structured streaming is a specific Spark library for handling live data streams, not a monthly billing pattern.
Compared with batch processing, what is a key characteristic of stream processing's data scope?
- AIt stores results only in a relational data warehouse
- BIt can process the entire dataset in a single scheduled operation, typically overnight
- CIt typically only has access to the most recent data, such as a rolling window✓
- DIt requires all input data to be fully prepared and validated before processing begins
Answer: Stream processing typically works with only the most recent data received, such as the last thirty seconds, rather than the whole dataset. Processing the entire dataset in one scheduled operation instead describes batch processing's broader data scope. Requiring input data to be fully prepared beforehand is also a batch processing trait, and stream processing results can be written to many sink types, not only a relational warehouse.
A building monitoring system must trigger alarms and unlock doors immediately when smoke and heat are detected. Which processing approach best fits this time-critical scenario?
- ABatch processing
- BStream processing✓
- CLambda architecture
- DELT processing
Answer: Stream processing handles time-critical operations that need an instant real-time response, such as triggering alarms the moment a sensor detects smoke or heat. Batch processing collects data before processing it together, which would introduce an unacceptable delay in an emergency. A lambda architecture combines batch and stream processing at a solution level rather than describing a single response mechanism, and ELT processing is a data-loading order for analytics, not a real-time alerting method.
In the lambda architecture pattern for combining batch and stream processing, what happens to streaming data when real-time analytics isn't required?
- AIt's written to the data store for subsequent batch processing✓
- BIt's sent directly to Power BI for live visualization
- CIt's processed only through Spark Structured Streaming and never stored
- DIt's discarded immediately after the event is generated and logged
Answer: When real-time analytics isn't needed, captured streaming data is instead written to the data store, similar to redirecting cars into a parking lot before counting them, so it can be handled later by batch processing. Discarding the data would prevent any later historical analysis, which the lambda pattern is designed to support. The data isn't restricted to Spark Structured Streaming processing alone, and sending it directly to Power BI describes real-time visualization, which is the scenario where analytics is required instead.
Which alternative to the lambda architecture eliminates the separate batch layer entirely, treating all data as a continuous stream and replaying it when historical reprocessing is needed?
- AKappa architecture✓
- BDelta architecture
- CA data lakehouse
- DA star schema
Answer: The kappa architecture is a simpler alternative to lambda that removes the distinct batch layer, treating everything as a continuous stream that can be replayed for historical reprocessing. A delta architecture is another combined batch-and-stream pattern but, unlike kappa, isn't described as eliminating the batch layer. A star schema is a data warehouse table design, and a data lakehouse is a hybrid storage architecture — neither is a processing-architecture pattern.
In a general stream processing architecture, what role does an output or sink play?
- AIt's the queue that guarantees each event is processed only once, in order, without duplication
- BIt's where the results of stream processing are written, such as a file or dashboard✓
- CIt's the perpetual query that selects, projects, and aggregates event data over time continuously
- DIt's the source that generates the original event data, such as a sensor device out in the field
Answer: The sink is where processed results are written, which could be a file, a database table, a real-time dashboard, or another queue for downstream processing. The original event generator, such as a sensor, is the source of the data rather than the destination. A queue that guarantees ordered, once-only processing describes a streaming source, and the perpetual query is the processing step that happens before results reach the sink.
Which Azure service ingests high volumes of event data, delivers events within a partition in order, and guarantees at-least-once delivery?
- AAzure Event Hubs✓
- BAzure Data Lake Store Gen 2
- CApache Kafka
- DAzure IoT Hub
Answer: Azure Event Hubs ingests high volumes of event data, delivers events within a partition in order, and guarantees at-least-once delivery. Azure IoT Hub is similar but is optimized specifically for managing event data from Internet-of-things devices. Azure Data Lake Store Gen 2 is a scalable storage service more often used for batch processing, and Apache Kafka is an open-source ingestion solution commonly paired with Apache Spark rather than an Azure-native service.
Which Microsoft Fabric Real-Time Intelligence component is a database optimized for time-series and event data, queried using KQL?
- AEventstream
- BActivator
- CEventhouse✓
- DReal-Time Dashboards
Answer: Eventhouse is the database in Fabric Real-Time Intelligence optimized for time-series and event data, queried using Kusto Query Language. Eventstream instead ingests, routes, and transforms streaming data before it reaches a destination such as Eventhouse. Activator triggers automated actions when streaming data meets defined conditions, and Real-Time Dashboards are used for live data visualization rather than storage and querying.
A team wants to define streaming jobs that ingest data from a source, apply a perpetual query, and write results to an output, outside of a Microsoft Fabric workspace. Which Azure service is described as a solid option for this standalone or hybrid scenario?
- ASpark Structured Streaming
- BAzure Stream Analytics✓
- CAzure Event Hubs
- DMicrosoft Fabric Real-Time Intelligence
Answer: Azure Stream Analytics is a platform-as-a-service solution for defining streaming jobs that ingest from a source, apply a perpetual query, and write results to an output, and it's described as a solid choice for standalone or hybrid scenarios outside of Fabric. Microsoft Fabric Real-Time Intelligence provides similar capability but is built specifically into a Fabric workspace. Spark Structured Streaming is a library for Spark-based services rather than a standalone PaaS job service, and Azure Event Hubs is an ingestion source rather than a stream-processing service.
Which open-source library enables Apache Spark to treat a live data stream as a continuously growing table?
- ALakeflow Declarative Pipelines
- BFabric Eventstream
- CDelta Lake
- DSpark Structured Streaming✓
Answer: Spark Structured Streaming lets you work with a live data stream the same way you'd work with a table, except the table keeps growing in real time as new data arrives. Delta Lake instead is an open-source storage format that adds reliability and schema enforcement to files in a data lake. Lakeflow Declarative Pipelines is an Azure Databricks ingestion framework, and Fabric Eventstream is a Microsoft Fabric component for routing streaming events rather than a Spark library.
Which capability of Delta Lake allows the same table to serve as both a real-time streaming sink and a source for historical batch queries?
- AReliability through change tracking
- BUnified batch and streaming✓
- CDirect Lake mode for querying OneLake files directly
- DSchema enforcement
Answer: Delta Lake's unified batch and streaming capability means the same table can be written to in real time as a streaming sink while also being queried as a source for historical batch analysis, avoiding separate storage. Schema enforcement instead ensures that data matches a defined structure before it's accepted, which is a data-quality guarantee rather than a dual-purpose storage capability. Reliability through change tracking prevents partial or failed writes from corrupting data, and Direct Lake mode is a Power BI query mode, not a Delta Lake table capability.