A call-level sentiment score tells you a conversation went badly. It does not tell a model where, and a model that cannot locate the moment cannot warn a supervisor while the call is still live.
Labeling each speaker turn gives escalation and agent-assist models something to fire on, and gives QA teams a way to audit a review rather than take a single number on faith. Every batch is labeled by several annotators and shipped with per-class inter-annotator agreement, so QA can tell which escalation calls the model is confident about and which ones still need a human read.