AWS Certified AI Practitioner
Evaluation Metrics: ROUGE, BLEU, and BERTScore
The three metrics the exam names, worked with real sentences and real numbers, plus the metrics AWS actually computes for each task type and the blind spot all of them share.
Intermediate 20 minutes 5 Learning Objectives
- Define ROUGE and compute ROUGE-1 and ROUGE-L on a real sentence pair
- Explain how BLEU combines n-gram precision with a brevity penalty and why it is a corpus-level metric
- Explain how BERTScore differs from word-overlap metrics and what it catches that they miss
- Match each metric to the task it suits and name the blind spot they all share
- Identify which computed metric AWS uses for each evaluation task type
