Text Analysis & NLP Concepts
10 free practice questions with explanations
10 free questions · instant explanations · no sign-up
PassNova has 10 free Microsoft AI-901 (Azure AI Fundamentals) practice questions on Text Analysis & NLP Concepts, each with a clear explanation. Practise them in the browser with instant feedback — 100% free, no sign-up, on any device. Updated for 2026.
Text Analysis & NLP Concepts: example questions & answers
10 worked examples with answers and explanations below. Practise them in the browser with instant feedback on every answer.
A data analyst is about to analyse a large body of text, referred to as a corpus. According to the AI-901 learning path, what is the opening step of that analysis?
- ACounting the frequency of each term
- BEncoding each word as an embedding vector
- CRemoving stop words such as 'the' and 'a'
- DBreaking the text down into tokens✓
Answer: Tokenisation breaks a corpus down into tokens, which are usually individual words but can also be partial words or combinations of words and punctuation. Each token is assigned a discrete numeric identifier, and only then can techniques such as frequency counting or embedding be applied to the text. Stop word removal is an optional refinement that excludes low-meaning words from the analysis rather than a step that precedes tokenisation.
A text analysis pipeline reduces 'running' to 'run' and 'global' to 'globe', applying linguistic rules and vocabulary so that the result is always a valid dictionary word. Which pre-processing technique is being used?
- AN-gram extraction
- BLemmatisation✓
- CStemming
- DText normalisation
Answer: Lemmatisation reduces words to their base or dictionary form, called a lemma, using linguistic rules and vocabulary so the output is a valid word. Stemming also consolidates related words but simply chops off endings such as 's', 'ing' and 'ed', which can leave a root that is not a real word. N-gram extraction finds multi-word phrases, and text normalisation covers steps such as removing punctuation and lower-casing before tokens are generated.
In the terminology used by the AI-901 learning path, what is a two-word phrase such as 'natural language' called when it is extracted from text as a single unit?
- AA bigram✓
- BA unigram
- CA trigram
- DA lemma
Answer: A bigram is a two-word phrase, one of the multi-term sequences found by N-gram extraction. A unigram is a single word and a trigram is a three-word phrase, so neither fits a two-word term. A lemma is the base dictionary form of a word produced by lemmatisation, not a measure of phrase length.
Which text analysis task identifies the people, organisations and locations mentioned in a document?
- ANamed entity recognition✓
- BParts of speech tagging
- CKeyword extraction
- DSentiment analysis
Answer: Named entity recognition identifies entities such as people, organisations and locations, and semantic models can be fine-tuned to examine each token's embedding and context to decide whether it is an entity and of what type. Parts of speech tagging labels tokens with grammatical categories such as noun or verb rather than real-world entities. Sentiment analysis classifies text by emotional tone, and keyword extraction finds the terms that best represent a document's main topics.
Two documents in a corpus of AI documentation both use 'agent', 'Microsoft' and 'AI' most often, so simple word counts cannot tell them apart. Which technique scores a term by how often it appears in one document compared with how common it is across the whole collection?
- ATF-IDF scoring✓
- BBag-of-words
- CFrequency counting
- DTextRank ranking
Answer: Term Frequency - Inverse Document Frequency (TF-IDF) assigns a high score to a word that appears often in one document but rarely across the rest of the corpus, giving it discriminative weight. Plain frequency counting only ranks terms within a single document and is exactly what failed to separate the two samples. Bag-of-words is a feature-extraction representation used as input to classifiers, and TextRank is a graph-based algorithm used mainly for summarisation and keyword extraction.
Which formula does the learning path give for the TF-IDF score of a term t in a document d, where N is the total number of documents and df(t) is the number of documents containing t?
- Atfidf(t, d) = tf(t, d) * log(N / df(t))✓
- Btfidf(t, d) = tf(t, d) + log(N / df(t))
- Ctfidf(t, d) = tf(t, d) * log(df(t) / N)
- Dtfidf(t, d) = log(tf(t, d)) * (N / df(t))
Answer: TF-IDF multiplies a term's frequency in a document by its inverse document frequency, idf(t) = log(N / df(t)). A term that appears in every document has df(t) = N, giving log(1) = 0 and no discriminative weight, which is why 'agent', 'Microsoft' and 'AI' scored zero across the two AI-agent samples in the learning path. Adding the two factors, inverting the ratio inside the logarithm or dividing by term frequency would all break this behaviour.
A developer wants to train a spam filter that predicts whether an email is spam from how often words such as 'miracle cure' and 'lose weight fast' occur, ignoring grammar and word order. Which feature extraction technique and classifier does the learning path describe for this?
- ATextRank graph with iterative rank scoring
- BContextualised embeddings with a transformer network
- CBag-of-words features with a Naive Bayes classifier✓
- DWord embeddings with a cosine-similarity classifier
Answer: Bag-of-words represents text as a vector of word frequencies or occurrences, ignoring grammar and word order. Naive Bayes is a probabilistic classifier that applies Bayes' theorem to those frequencies to predict a document's class, and the same pairing can assign sentiment labels such as positive or negative. Embedding-based and transformer approaches capture contextual meaning rather than raw counts, and TextRank is a graph algorithm for summarisation and keyword extraction rather than a classifier.
A summarisation tool treats each sentence as a node in a graph, weights the edges by word overlap between sentences, ranks the nodes iteratively and then returns the highest-ranked sentences unchanged. Which approach is this?
- AText classification with bag-of-words
- BKeyword extraction with TF-IDF scores
- CExtractive summarisation with TextRank✓
- DAbstractive summarisation with a semantic model
Answer: TextRank is an unsupervised graph-based algorithm that applies the same principle as Google's PageRank to text, with a damping factor typically set to 0.85. Because the summary consists of a subset of the original sentences and no new text is generated, the result is extractive summarisation. Abstractive summarisation generates new language to express the key themes, while keyword extraction and text classification produce terms or category labels rather than a summary.
Using three-dimensional embeddings, a developer calculates a cosine similarity of 0.992 between 'dog' and 'cat' but only 0.333 between 'dog' and 'tree'. What does the high value for 'dog' and 'cat' indicate?
- ATheir tokens share the same numeric identifier in the vocabulary
- BTheir vectors point in similar directions, reflecting related meaning✓
- CTheir words co-occur in one document more often than 'tree' does
- DTheir vectors have equal magnitudes, reflecting equal frequency
Answer: Cosine similarity compares the orientation of two vectors, and semantically similar tokens produce embeddings that point in similar directions in multidimensional space. A score near 1 therefore means the words are used in similar contexts and carry related meaning, while 'tree' has a distinctly different orientation. The measure is independent of vector magnitude, has nothing to do with token identifiers, and does not count co-occurrence within a document.
Using the learning path's example embeddings, a developer calculates kitten - puppy + dog and searches for the word whose vector is closest to the result. Which word is found, and what kind of reasoning does this demonstrate?
- Ayoung, demonstrating vector translation
- Btree, demonstrating odd-one-out detection
- Ccat, demonstrating analogical reasoning✓
- Ddog, demonstrating cosine similarity
Answer: Vector arithmetic on embeddings can answer analogy questions such as 'puppy is to dog as kitten is to ?'. Subtracting puppy from kitten and adding dog gives [0.7, 0.5, 0.2], which matches the vector for cat, so the operation performs analogical reasoning. Vector translation is the simpler addition or subtraction of a single word such as 'young', cosine similarity measures orientation rather than solving analogies, and odd-one-out detection compares pairwise similarities.