Been burned by exactly the problem this solves - manually testing voice models against a handful of sample clips every time a new one drops, then forgetting to re-test six months later when something better shipped. Having routing decided against real public benchmarks instead of vendor marketing numbers is the actual value here, not just "one API for everything."