Skip to article frontmatterSkip to article content
Site not loading correctly?

This may be due to an incorrect BASE_URL configuration. See the MyST Documentation for reference.

Design goals

In short, the aim of speech coding methods is primarily to enable natural and efficient spoken communication over a geographical distance, given constraints on available resources. In other words, we want to be able to talk with a distant person with the aid of technology. Usually distance refers to location, but speech coding can be (and is often) used for storing speech signals (such that distance refers to distance in time). Bäckström et al., 2017

In particular, aspects of quality which can be included in our design goals are for example:

It is clear that different types of quality are prominent at different levels of coding-accuracy, which in turn is a function of available bitrate (bandwidth). As a rough characterization, with classic signal processing (non-neural) codecs, the quality-issues we optimize at different bitrates are:

With neural speech codecs, the quality trade-offs are somewhat different Zeghidour et al., 2021Défossez et al., 2022Kumar et al., 2023:

Performance of a codec is however always a compromise between quality and resources. By increasing the amount of computational resources (or bandwidth) we can improve quality ad infinitum. The most important limited resources are

Furthermore, the use-case of the intended speech (and audio) codec has many important effects on the overall design. For example, the systems configuration can be one of the following:

The overall design is also influenced by the type of transmission link. In particular, the first few generations of digital mobile phones operated with circuit-switched networks, where a fixed amount of bandwidth is allocated to every connection. Newer networks are however based on packet-switched designs, where data is transmitted essentially over the internet and capacity and routing is optimized on the fly. Packet-switched networks can in practice be much better optimized for overall cost and performance. However, a packet-switched network cannot guarantee a steady flow of packets, such that the receiver has to tolerate both delayed or missing packets as well as packets which arrive in the wrong order. Clearly this has an impact on both overall transmission delay of the system, as well as increases the computational complexity of the receiver. The costs are however usually balanced by the savings gained in network optimization.

A further important related aspect are assumptions about lost packets in general. In many storage and broadcast applications we can assume that packets are not lost and that all data is available at the receiver. It however much more common that we must assume that some packets are lost. Among the most important consequences of lost packets for the design is that in decoding the signal, we cannot assume that we have access to previous packets. Specifically, if decoding of the current packet depends on the previous packet, then a single lost packet would make us unable to decode any of the following packets. Clearly such sensitivity to lost packets is unacceptable in most real-world transmission systems. However, we could encode speech with much higher efficiency, if we were allowed to use previous packets to predict the current packet. The likelihood of lost packets thus dictates the compromise between sensitivity to lost packets and coding (compression) efficiency.

Observe that the usefulness of very low rates, below roughly 5-7 kbit/s, is narrower than the 7-32 kbit/s range. The bottleneck is caused by the fixed overheads in the transmission layer, where every packet carries regardless of its payload size: a typical RTP/UDP/IP header stack adds around 40 bytes per packet for IPv4 (RTP 12 bytes + UDP 8 bytes + IP 20 bytes), before any link-layer framing is even added. With 20 ms steps between windows, we have 50 windows per second; if each packet carries only one window, a codec running at 5.9 kbit/s contributes only around 15 bytes of payload per packet - so roughly 70% of every packet is header, not speech data. Lowering the codec bitrate further barely reduces the bandwidth actually used, since the packet is already dominated by fixed overhead, and the marginal benefit of further compression all but disappears.

The only way to recover further bandwidth savings below this point is to amortize the header cost over several frames per packet, by increasing the delay between transmissions - the same delay-versus-efficiency tradeoff as above, just applied at the low-bitrate end rather than at the packet-size limit. This pattern appears in other low-rate speech RTP formats too: RFC 3558 (EVRC/SMV) and RFC 4788, for instance, define bundled or interleaved packet formats specifically to amortize the RTP header over more than one speech frame.

Rates below roughly 5-6 kbit/s are thus not very useful in everyday, low-delay use cases without such frame bundling, but remain relevant in special applications that can tolerate the extra delay or where bandwidth is severely constrained, such as military and satellite communication, as well as storage. Another potential application is cases where some other information is transmitted in the same packets, such as multi-channel audio or multimedia. The sanity can however again be contested, as the single-channel quality might not be sufficiently high to warrant multi-channel audio, and transmission of video requires several orders of magnitude larger capacity anyway, so greedy optimization of speech compression brings only very marginal benefit in the total budget.

References

References
  1. Bäckström, T., Lecomte, J., Fuchs, G., Disch, S., & Uhle, C. (2017). Speech coding: with code-excited linear prediction. Springer. 10.1007/978-3-319-50204-5
  2. Zeghidour, N., Luebs, A., Omran, A., Skoglund, J., & Tagliasacchi, M. (2021). SoundStream: An End-to-End Neural Audio Codec. IEEE/ACM Transactions on Audio, Speech, and Language Processing, 30, 495–507. 10.1109/TASLP.2021.3129994
  3. Défossez, A., Copet, J., Synnaeve, G., & Adi, Y. (2022). High Fidelity Neural Audio Compression. 10.48550/ARXIV.2210.13438
  4. Kumar, R., Seetharaman, P., Luebs, A., Kumar, I., & Kumar, K. (2023). High-Fidelity Audio Compression with Improved RVQGAN. 10.48550/ARXIV.2306.06546