Microsoft DP-900 (Azure Data Fundamentals)

Data Formats, Storage & Databases

12 free practice questions with explanations

12 free questions · instant explanations · no sign-up

PassNova has 12 free Microsoft DP-900 (Azure Data Fundamentals) practice questions on Data Formats, Storage & Databases, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.

Sample questions

Data Formats, Storage & Databases: example questions & answers

12 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.

  1. A dataset in which every record has the same fields, typically organized into rows and columns, is described as which type of data?

    • AStructured data✓
    • BSemi-structured data
    • CUnstructured data
    • DVector data

    Answer: Structured data adheres to a fixed schema, so every instance of an entity has the same fields or properties, most often arranged as rows and columns in a table. Semi-structured data allows some variation between instances, such as optional fields. Unstructured data, like images or video, has no defined schema at all. Vector data represents embeddings used by AI systems rather than tabular records.

  2. Two JSON documents both describe a customer, but one includes two phone numbers under 'contact' while the other includes none. This flexibility, where the schema can vary between instances, is characteristic of which kind of data?

    • ASemi-structured data✓
    • BRelational data
    • CCompletely unstructured data
    • DFully structured data

    Answer: Semi-structured data has some structure but allows variation between entity instances, such as optional or repeating fields, and JSON is a common format for representing it. Structured data requires every instance to share exactly the same fixed fields. Unstructured data, such as documents or audio files, has no defined schema. Relational data is structured data organized into tables linked by keys.

  3. Which of these is an example of unstructured data?

    • AA JSON document with fixed attributes
    • BA video file✓
    • CA comma-separated values file
    • DA table of customer records

    Answer: Video, along with images, audio, and many binary documents, has no specific structure and is classed as unstructured data. A comma-separated values file and a table of customer records both organize data into a consistent schema, making them structured. A JSON document with fixed attributes still follows a defined shape, so it is not unstructured.

  4. What is a common format for storing data as plain text, with fields separated by commas and rows ended by a new line?

    • AApache Parquet columnar storage
    • BExtensible Markup Language (XML)
    • CBinary Large Object (BLOB) storage
    • DComma-separated values (CSV)✓

    Answer: Comma-separated values (CSV) is the most common delimited text format, with commas separating fields and a new line marking the end of each row. Extensible Markup Language (XML) instead represents data using angle-bracket tags. Binary Large Object (BLOB) storage holds raw binary bytes rather than delimited text. Apache Parquet columnar storage is a compressed, structured format, not a plain-text one.

  5. Which data format represents entities as hierarchical documents built from objects enclosed in braces and collections enclosed in square brackets?

    • AExtensible Markup Language (XML)
    • BComma-separated values (CSV)
    • CJavaScript Object Notation (JSON)✓
    • DApache Avro

    Answer: JavaScript Object Notation (JSON) defines data entities as objects enclosed in braces, with collections enclosed in square brackets, making it a flexible hierarchical format. Extensible Markup Language (XML) instead represents elements using angle-bracket tags. Comma-separated values (CSV) is a flat, delimited text format rather than a hierarchical one. Apache Avro stores its structure in a JSON header but encodes the actual records as binary, not as readable braces and brackets.

  6. A human-readable format that uses tags enclosed in angle brackets to define elements, and that has largely been superseded by JSON, is called what?

    • AJavaScript Object Notation (JSON)
    • BComma-separated values (CSV)
    • CExtensible Markup Language (XML)✓
    • DApache Parquet

    Answer: Extensible Markup Language (XML) uses tags enclosed in angle brackets and was popular in the 1990s and 2000s, but it has largely been superseded by the less verbose JSON format. JavaScript Object Notation (JSON) is the format that superseded it, not the older tag-based one. Comma-separated values (CSV) separates fields with commas rather than tags. Apache Parquet is a modern columnar binary format unrelated to angle-bracket tags.

  7. Which optimized file format is columnar, stores data for each column together within row groups, and is described as the de facto standard for modern data lakehouses?

    • ADelta Lake
    • BExtensible Markup Language (XML)
    • CApache Avro
    • DApache Parquet✓

    Answer: Apache Parquet is a columnar format whose files contain row groups, with data for each column stored together and described by metadata, making it efficient for lakehouse storage and processing. Apache Avro is row-based rather than columnar, storing its structure in a JSON header. Delta Lake builds on Parquet by adding a transaction log rather than being a columnar format itself. Extensible Markup Language (XML) is a text-based markup format, not a compressed columnar one.

  8. Which optimized file format stores a header describing the data's structure as JSON, then stores the actual records as binary data in one or more blocks?

    • AApache Parquet
    • BBinary Large Object (BLOB)
    • CApache Avro✓
    • DDelta Lake

    Answer: Apache Avro files contain a header that describes the structure of the data as JSON, with the records themselves stored as binary information in one or more blocks. Apache Parquet is columnar and organizes data into row groups rather than header-described blocks. Delta Lake adds a transaction log on top of Parquet files rather than defining its own header-and-block layout. A Binary Large Object (BLOB) is raw binary data with no structural header at all.

  9. Which open-source storage format builds on Parquet by adding a transaction log that enables ACID transactions, data versioning, and reliable updates in a data lake?

    • AComma-separated values (CSV)
    • BDelta Lake✓
    • CApache Avro
    • DExtensible Markup Language (XML)

    Answer: Delta Lake builds on Parquet by adding a transaction log, which enables ACID transactions, data versioning, and reliable updates for files stored in a data lake. Apache Avro is a separate row-based format with no transaction log of its own. Comma-separated values (CSV) is a plain delimited text format with no transactional capability. Extensible Markup Language (XML) is an older tag-based markup format, also without a transaction log.

  10. Data professionals commonly refer to files of raw binary data, such as images or video, that must be interpreted and rendered by an application, using which term?

    • AComma-separated values (CSV)
    • BExtensible Markup Language (XML)
    • CBinary Large Object (BLOB)✓
    • DApache Parquet

    Answer: Files stored as raw binary data, including images, video, audio, and application-specific documents, are commonly referred to as Binary Large Objects, or BLOBs. Comma-separated values (CSV) and Extensible Markup Language (XML) both map bytes to printable characters through a character encoding scheme, so they are human-readable rather than raw binary. Apache Parquet is an optimized columnar format designed for structured data, not a term for raw binary files.

  11. In a relational database, using a customer's primary key to reference that customer from a sales order record, rather than repeating the customer's details in every order, is an example of what?

    • ADenormalization
    • BSharding
    • CPartitioning
    • DNormalization✓

    Answer: Normalization uses key values to reference data entities in other tables, eliminating duplicate data so that a customer's details are stored only once rather than repeated for every sales order. Denormalization deliberately reintroduces some duplication, typically to make analytical queries perform faster. Partitioning and sharding both describe ways of splitting data across storage, not of removing duplication through keys.

  12. A nonrelational database that stores tabular data but groups its columns into logically related sets, rather than applying a single relational schema, is known as which type?

    • AGraph database
    • BColumn family database✓
    • CDocument database
    • DKey-value database

    Answer: A column family database stores tabular data comprising rows and columns, but divides the columns into groups known as column families that hold logically related sets of columns. A document database is a form of key-value database in which the value is a JSON document. A key-value database pairs each unique key with an associated value in any format. A graph database stores entities as nodes with links that define relationships between them.

Start practising Data Formats, Storage & Databases →