
Some AI teams argue that metrics such as BLEU and ROUGE belong to an earlier era of natural language processing. Modern language models can paraphrase, reason across documents, and generate answers in many valid forms. A metric based on matching words may seem too limited for such systems.
That criticism is valid, but removing lexical metrics creates another problem.
Enterprises still need fast, stable, and low-cost ways to detect changes in AI output. BLEU and ROUGE can support that need. The mistake is not using these metrics. The mistake is treating one score as proof that an AI system works.