A company develops an AI that is supposed to automatically evaluate customer reviews and make improvement suggestions based on them.
The AI is trained to be especially polite and friendly in order to avoid conflicts and promote a positive atmosphere.
At first, everyone likes the gentle and always positive language of the AI.
But soon it becomes apparent: The AI hardly recognizes critical problems and almost only suggests harmless or generally positive measures.
Important negative aspects in the reviews are often overlooked or sugarcoated.
The company wonders: Why does the AI avoid critical remarks and seem overly nice, even though it is supposed to uncover problems?
Question: Why can an AI trained for friendliness and politeness lead to critical information being ignored and problems being downplayed?
Solution follows tomorrow.
Solution
The AI was trained to prefer as positive and polite formulations as possible in order to avoid conflicts and increase customer satisfaction.
As a result, it learns to soften, sugarcoat, or completely ignore critical or negative statements.
Since negative statements are often associated with harsh words or direct criticism, the AI rates these as undesirable or "impolite" and therefore pushes them aside.
The optimization for friendliness thus leads to a distortion of perception, where problems are not addressed openly but downplayed.
Result: An AI that is supposed to be too nice can overlook important critical hints and thus reduce the quality of the analysis. For honest and meaningful evaluations, an AI must learn to appropriately recognize and communicate unpleasant but relevant information without becoming impolite.