Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Objective quality evaluation

Objective estimators for perceptual quality

With “objective evaluation” we usually refer to estimators of perceptual quality, where the objective is to predict the mean output of a subjective listening test using an algorithm. That is, we want a computer to listen to a sound sample and try to “guess” what a human listener would say about its quality (on average).

It is then clear that subjective evaluation is always the “true” measure of performance, and objective evaluation is an approximation thereof. In this sense, subjective evaluation is “better”. There are plenty of examples where objective quality estimators give the opposite result of the subjective preference Manocha et al., 2022. However, there are many good reasons to use objective instead of subjective evaluation:

Some of the most frequently used objective measures include:

Other objective performance criteria

There are many cases where other performance criteria are well-warranted than mere prediction of subjective listening test results. Most typically, these criteria are applied when there is no user involved, such as speech recognition or, when we want to have more detailed characterization of performance than given by predictors of subjective listening test results.

Some examples of such performance criteria include:

References
  1. Manocha, P., Jin, Z., & Finkelstein, A. (2022). Audio Similarity is Unreliable as a Proxy for Audio Quality. arXiv Preprint arXiv:2206.13411. 10.48550/arXiv.2206.13411
  2. Rix, A. W., Beerends, J. G., Hollier, M. P., & Hekstra, A. P. (2001). Perceptual evaluation of speech quality (PESQ)-a new method for speech quality assessment of telephone networks and codecs. 2001 IEEE International Conference on Acoustics, Speech, and Signal Processing. Proceedings (Cat. No. 01CH37221), 2, 749–752. 10.1109/ICASSP.2001.941023
  3. Beerends, J. G., Schmidmer, C., Berger, J., Obermann, M., Ullmann, R., Pomy, J., & Keyhl, M. (2013). Perceptual objective listening quality assessment (POLQA), the third generation ITU-T standard for end-to-end speech quality measurement part i—temporal alignment. Journal of the Audio Engineering Society, 61(6), 366–384. http://www.aes.org/e-lib/browse.cfm?elib=16829
  4. Thiede, T., Treurniet, W. C., Bitto, R., Schmidmer, C., Sporer, T., Beerends, J. G., & Colomes, C. (2000). PEAQ - The ITU standard for objective measurement of perceived audio quality. Journal of the Audio Engineering Society, 48(1/2), 3–29. http://www.aes.org/e-lib/browse.cfm?elib=12078
  5. Taal, C. H., Hendriks, R. C., Heusdens, R., & Jensen, J. (2011). An algorithm for intelligibility prediction of time – frequency weighted noisy speech. IEEE Transactions on Audio, Speech, and Language Processing, 19(7), 2125–2136. 10.1109/TASL.2011.2114881
  6. Gray, A., & Markel, J. (1976). Distance measures for speech processing. IEEE Transactions on Acoustics, Speech, and Signal Processing, 24(5), 380–391. 10.1109/TASSP.1976.1162849