Why tolerate bad phone audio?

I also cannot see the point of having HD quality telephone connections, perhaps someone can enlighten me as to what they are good for. Perhaps those telemarketing people, or customer service folks would sound more intelligent and actually be able to correctly diagnose and fix a simple problem once and for all. If HD quality does that then I am game otherwise skip this nonsense and bring on video calling so the person on the other side can see the look of frustration on my face when I have to put with their crap.

While the world has moved on since the 1930's with improvements in TV, sound recording etc, voice has stayed reasonably static. With High Definition voice, you finally get the quality in line with the 21st century. It's a case of 'hear for yourself' - once you hear it you get hooked. It leads to richer voice tones (good with family conversations), ease of understanding, easier listening in phone teleconferences etc. It paves the way for voice-understanding with IVR's (now still very frustrating). The codecs used can also be more resilient and bandwidth efficient in the presence of line noice.

This is not so much a case of 'marketing' as the industry catching up. Fortunately France Telecom, Orange and British Telecom are paving the way. In the United States and Britain, many of the smaller players are paving the way - hence the need for the neutral sip peering hub to help people talk across networks and so build up critical mass
 
I wouldn't go to such extremes to bash
I would. This article is another crock of misleading marketing.

Human speech safely ends at about 10 to 12 kHz (baseband ofcourse), the amount of power our vocal cords can produce above those frequencies is minimal if not negligible.
I'm very curious to know which humans on this planet can reach anywhere close to that. If you've seen 5th Element, you might recall the Diva Dance scene performed by a professional soprano vocalist whose voice had to be synthesized for some of the notes because it couldn't reach high enough, and those notes were well below 10 kHz. Vocal pitches above 3 kHz are attainable only by women with world records to their name.

Hopefully astute readers will realise that if a system designed to carry human voice constituting 20% of our audible range is changed to carry 100% of our audible range, that's a potential for it to carry 80% of noise. Better audibility? Don't count on it. A good money/PR spinner? Definitely.

The only thing telephones really need to carry is 100% of our vocal range, not 100% of our audible range.
 
I would. This article is another crock of misleading marketing.


I'm very curious to know which humans on this planet can reach anywhere close to that. If you've seen 5th Element, you might recall the Diva Dance scene performed by a professional soprano vocalist whose voice had to be synthesized for some of the notes because it couldn't reach high enough, and those notes were well below 10 kHz. Vocal pitches above 3 kHz are attainable only by women with world records to their name.

Hopefully astute readers will realise that if a system designed to carry human voice constituting 20% of our audible range is changed to carry 100% of our audible range, that's a potential for it to carry 80% of noise. Better audibility? Don't count on it. A good money/PR spinner? Definitely.

The only thing telephones really need to carry is 100% of our vocal range, not 100% of our audible range.

Well there you go (I don't do the biological side of things :p). But either way, the engineering/science content of this article was dismally incorrect just that by usually claiming 'idiots' and other similar terms gets your point labelled as moot and a personal attack.
 
Last edited:
Som much has been achieved in the telephony space this year. Adoption of HD Voice in SA is a wonderful opportunity for us to rank with the top in the world and prove South Africa's innovative and entrepreneurial spirit'

After reading (widely available) articles on High Definition voice, I think the 70 000 Hz mentioned is a typo and should read 7 000 Hz. That fits with what other people report and it makes sense because the article does go on to say 'double the audio bandwidth that most of us experience today' (ie with voice calls).

In other words, HD Voice will extend the bass range down to 80 Hz. This is compared with the current bass limit for audio codecs of 300 Hz, which is very high for speech (actually above the Middle C note at 262 Hz - meaning you can picture half a piano not being audible!). A lower limit of 80Hz captures lower voice frequencies and I am sure it is based on research on what voice frequencies enhance the communication experience.

The upper frequency limit of current voice is 3.4 kHz. The range of 3.4kHz less 3.4kHz gives a 3.1kHz range, or 20% of what most humans can hear. If HD Voice typically captures up to 7 kHz (not 70!), then it is just over twice the frequency now being used.

If HD can achieve this wider band, still working within the 'useful' voice range, but now 'more useful' and do it with more efficient codecs, then let's look beyond a single typo in an article and look at the bigger picture.
 
Last edited:
Som much has been achieved in the telephony space this year. Adoption of HD Voice in SA is a wonderful opportunity for us to rank with the top in the world and prove South Africa's innovative and entrepreneurial spirit'

After reading (widely available) articles on High Definition voice, I think the 70 000 Hz mentioned is a typo and should read 7 000 Hz. That fits with what other people report and it makes sense because the article does go on to say 'double the audio bandwidth that most of us experience today' (ie with voice calls).

In other words, HD Voice will extend the bass range down to 80 Hz. This is compared with the current bass limit for audio codecs of 300 Hz, which is very high for speech (actually above the Middle C note at 262 Hz - meaning you can picture half a piano not being audible!). A lower limit of 80Hz captures lower voice frequencies and I am sure it is based on research on what voice frequencies enhance the communication experience.

The upper frequency limit of current voice is 3.4 kHz. The range of 3.4kHz less 3.4kHz gives a 3.1kHz range, or 20% of what most humans can hear. If HD Voice typically captures up to 7 kHz (not 70!), then it is just over twice the frequency now being used.

If HD can achieve this wider band, still working within the 'useful' voice range, but now 'more useful' and do it with more efficient codecs, then let's look beyond a single typo in an article and look at the bigger picture.

The ultimate problem with this concept is and always will be: why do I need to make my expensive GSM base-stations redundant, increase the load on my UMTS network considerably and have to invest in new base-stations everywhere because someone wants to listen to higher quality voice that will be distorted by the tiny speaker it's played on anyways?
 
The ultimate problem with this concept is and always will be: why do I need to make my expensive GSM base-stations redundant, increase the load on my UMTS network considerably and have to invest in new base-stations everywhere because someone wants to listen to higher quality voice that will be distorted by the tiny speaker it's played on anyways?

Yeah why do we need cars that produces less CO2, when the cars we have are perfectly good enough, rite? Or why do we need better roads when the existing pothole infested roads are good enough to get a car from A to B.

You don't seem to like the benefits of more advanced technology can do for us. The mentality of if it aren't broke, don't "fix" it doesn't help us to move forward.
 
Yeah why do we need cars that produces less CO2, when the cars we have are perfectly good enough, rite? Or why do we need better roads when the existing pothole infested roads are good enough to get a car from A to B.

You don't seem to like the benefits of more advanced technology can do for us. The mentality of if it aren't broke, don't "fix" it doesn't help us to move forward.

Lol, unlike those 2 examples you mentioned, 'HD voice' has few benefits.

My work is focused on advancing technology (wireless telecomms especially) where the point is to squeeze every last possible bit/s/Hz out of a channel, so I'd say it's pretty ironic accusing me of that.
 
Lol, unlike those 2 examples you mentioned, 'HD voice' has few benefits.

My work is focused on advancing technology (wireless telecomms especially) where the point is to squeeze every last possible bit/s/Hz out of a channel, so I'd say it's pretty ironic accusing me of that.

High Definition voice has a growing following based on customer demand. There are >500 000 France Telecom subscribers (as at 2009), British Telecom is rolling it out, Orange Mobile (on six networks),there are dozens of other networks picking up demand and a good proportion of Skype calls are now high definition. There is definitely demand.

This is one example of the wave of IP-based services customers want and which are better delivered using IP communications (especially SIP-based services, including Video, Instant Messaging, Unified Communications. These are all growing and unfortunately will make some of the traditional 'TDM' infrastructure obsolete. It is a new paradigm which empowers the innovative and entrepreneurial communications providers with IP networks.

What is holding back more growth is the restriction in making calls from one network to another, including different service providers, mobile-fixed and international. That is why we are promoting the vision of a SIP peering Hub, to enable customers to call across networks http://mybroadband.co.za/news/telecoms/16615-Boosting-VoIP-services.html

We hope to enable service providers to offer what customers really want and at better pricing. The revolution is advanced in some of our other markets (eg Korea, Netherlands, US and there is no reason local service providers cannot also be empowered.
 
High Definition voice has a growing following based on customer demand. There are >500 000 France Telecom subscribers (as at 2009), British Telecom is rolling it out, Orange Mobile (on six networks),there are dozens of other networks picking up demand and a good proportion of Skype calls are now high definition. There is definitely demand.

This is one example of the wave of IP-based services customers want and which are better delivered using IP communications (especially SIP-based services, including Video, Instant Messaging, Unified Communications. These are all growing and unfortunately will make some of the traditional 'TDM' infrastructure obsolete. It is a new paradigm which empowers the innovative and entrepreneurial communications providers with IP networks.

What is holding back more growth is the restriction in making calls from one network to another, including different service providers, mobile-fixed and international. That is why we are promoting the vision of a SIP peering Hub, to enable customers to call across networks http://mybroadband.co.za/news/telecoms/16615-Boosting-VoIP-services.html

We hope to enable service providers to offer what customers really want and at better pricing. The revolution is advanced in some of our other markets (eg Korea, Netherlands, US and there is no reason local service providers cannot also be empowered.

While all-IP networks and IMS/SIP solutions are indeed inevitable, the product still doesn't account for the omnipresence that is GSM. There's the catch-22 of GSM being so heavily invested that no company wishes to make it redundant and the fact that current GSM standards don't support the higher vocoder rates. This would mean that either the GSM standard would need to be changed (difficult, lots of money need to be spent) or stop using the GSM network and switch over to UMTS for your higher bandwidth requirements (a possible third solution would be to use EDGE, provided by the GSM networks, but you would need to develop some funky codecs [which, in my line of research, I have come to realise that it's only a matter of time and judging by the bitrates afforded, it should be able to handle]).

So it would boil down to: Try and somehow get 'HD Voice' in the GSM standard or make your GSM network obsolete and increase load on your UMTS network.

I'm pretty sure many providers are a bit reluctant with the latter since the whole point of the development of GSM's vocoders was to optimise 'understandability' of the encoded voice while keeping the bit-rate below a threshold.

In my opinion, unless someone actually does go with the EDGE idea, this whole service will remain opt-in (and in limited quantities) until we see all-IP networks completely overtaking/phasing out GSM.
 
While all-IP networks and IMS/SIP solutions are indeed inevitable, the product still doesn't account for the omnipresence that is GSM.

.....

In my opinion, unless someone actually does go with the EDGE idea, this whole service will remain opt-in (and in limited quantities) until we see all-IP networks completely overtaking/phasing out GSM.

Exactly.

Not only are the networks still heavily biased toward GSM, (guess around 3:1 in SA), but we need to take into account that GSM handsets are by far still dominant in SA with probably a much higher ratio to 3G handsets. And, in any case, 3G still controls the bandwidth allocated to the voice channel to save on precious resources.

Once we get to all-IP (most likely only with LTE), I assume you can use your allocation of bandwidth for whatever you want. But don't be surprised if the commercial models then change to usage-, and not time-, based.
 
I think it is clear that GSM is a dominant technology which has outlasted the predictions of many analysts and still has 'legs' into the future.

However I believe the GSM investment does not create inertia, and the adoption patterns with better voice will continue regardless.

One of the reasons for this has been its ability to adapt, whilst remaining backward compatible. We have seen first the rise of enhanced-full-rate (to improve voice), then half-rate (to cope with peak demand) and then AMR, or adaptive-modulation to cope with quality-vs-bandwidth tradeoffs.

Also base stations have changed. The cutting edge GSM base stations is actually a 2G/3G/LTE ready node with software defined radio and all-ip transmission. There is very little kit (as a fraction of the network investment) which is GSM exclusive and would constitute a protected investment.

In addition, most of the network investment in South Africa is new and mindful of 3G etc transitions. Entire provinces have already been swapped out within the last 2 years, to replace with more energy efficient, transmission efficient, better-radio quality systems. This is the case with all 3 GSM networks. Much of the core network is also all-IP.

On the adoption side, there are different scenarios:

1. In fixed networks, high definition voice will take off regardless. New handsets already support it as well as PBX's (Asterisk included). The ecosystem is in place, it is a matter of time. Operators are already deploying it. Only the 'interconnect constraint' will hold this back, and this is where we are pushing for a sip peering hub to expedite. We have already proven this with High Definition voice peering in the US and UK.
2. On mobile networks, there are three scenarios I can use to illustrate options:
2.1 the largest mobile operator goes high-definition, as a way to invest their excess cash and make their market dominance more unassailable
2.2 the smallest mobile operator goes high-def, to showcase their HSPA network, use their invested bandwidth and differentiate to high value customers
2.3 early adopters (normally those with cash and new handsets) run their own VoIP via SIP clients, show preference for high-def, and force operators to 'join the party'

In all these scenarios, the customer will only benefit if they have the freedome to call across networks to other high-def subscribers, creating a sustainable ecosystem large enough to justify the use of the new technology. That is why we are encouraging telcos to join our high definition voice initiative and have ready access to this growing base.
 
I'm very curious to know which humans on this planet can reach anywhere close to that. If you've seen 5th Element, you might recall the Diva Dance scene performed by a professional soprano vocalist whose voice had to be synthesized for some of the notes because it couldn't reach high enough ... Vocal pitches above 3 kHz are attainable only by women with world records to their name.

Hopefully astute readers will realise that if a system designed to carry human voice constituting 20% of our audible range is changed to carry 100% of our audible range, that's a potential for it to carry 80% of noise..

A common misconception! Astute readers with science 101 know that if you want to sound like a sine wave when you talk, then 3kHz is plenty. It's the harmonic frequencies which are higher frequences and give each voice it's personality and quality, from Pavarotti to Pavlova. This singer would have audible harmonics at 6,9,12,15, 18kHz. That's why audiophiles add tweeters to their sound systems. The opera singer might break records, but when singing over a standard telephone she would sound like a sine wave and she wouldn't sell any records!
 
Last edited:
Astute readers with science 101 know that if you want to sound like a sine wave when you talk, then 3kHz is plenty. It's the harmonic frequencies which are higher frequences and give each voice it's personality and quality, from Pavarotti to Pavlova. That's why people add tweeters to their sound systmes. The opera singer might break records, but when sung over a standard telephone it would sound like a sine wave and she wouldn't sell any records!

Lol, Pavlova was a ballet dancer, not a singer. Harmonics sure do make it, but past ~ 16 kHz, the average adult can't hear anything. Also, tweeters are there because it's difficult to move larger diaphragms at higher frequencies (inertia).
 
Lol, Pavlova was a ballet dancer, not a singer. Harmonics sure do make it, but past ~ 16 kHz, the average adult can't hear anything. Also, tweeters are there because it's difficult to move larger diaphragms at higher frequencies (inertia).

Just making the point that the harmonics are useful for better enjoyment of human speech over a phone. The current cut-off of 3.1kHz or so is rather low. I wouldn't actually want to listen to Pavarotti on the phone, but I would like to hear my kids or my spouse talk to me in faithful reproduction of their voices which I enjoy so much and I want to actually hear some of the mid-range harmonics that make it possible
 
Just making the point that the harmonics are useful for better enjoyment of human speech over a phone. The current cut-off of 3.1kHz or so is rather low. I wouldn't actually want to listen to Pavarotti on the phone, but I would like to hear my kids or my spouse talk to me in faithful reproduction of their voices which I enjoy so much and I want to actually hear some of the mid-range harmonics that make it possible

Sure thing, my argument still stands though that companies will be reluctant to throw away/invalidate their GSM networks (THE most ubiquitous of them all) for higher quality voice and on top of that add significantly more load to he UMTS networks. i.e. it won't be ubiquitous until GSM disappears. This, ofcourse, is because GSM was designed in a utilitarian view (i.e. if you can understand what the other person is saying, it's good enough) instead of an audiophile's point of view.

The other argument was that this article was very poorly written in terms of technical content, but that's unanimous.
 
Last edited:
The other argument was that this article was very poorly written in terms of technical content, but that's unanimous.

Well, one typo on the 70 000 Hz instead of 7000 Hz got people pretty worked up, without looking at the bigger picture.
 
As already explained by a number of posters above, the technicalities and drivers here are relatively simple and well understood:

- There is only so much spectrum available to each operator.
- Within a technology family (GSM, UMTS, HSPA, etc.) there is a finite amount of bandwidth available from this spectrum. Spectrum re-use, etc. also comes into play.
- Thus each tower has a finite capacity with which to service its footprint.
- High speed data requires lots of capacity.
- The less bandwidth you allocate per user (i.e. just enough to make a decent call), the more users you can serve concurrently.
- HD voice requires more bandwidth and thus you can serve fewer concurrent callers compared to SD voice.

In South Africa, we have constant challenges to increase capacity per base-station. In no particular order:

- Finite amount of spectrum.
- Extremely high SIM penetration, already over 100% and growing.
- Backhaul availability and capacity constraints.
- Failure of fixed line alternatives increase load on wireless solutions.
- Dominance of GSM, both in base-stations and handsets.
- Slow deployment and adoption of 3G (and 4G!) technologies for voice.
- Yet, fast uptake of 3G data services.
- Difficulty in getting more coverage (community resistance to rolling out new towers, etc).
- Price sensitivity of consumers (new networks are expensive and costs must be recovered).

It's clear pressure on existing capacity will only increase in the next few years. Thus, if anything, networks will look at REDUCING the bandwidth available per (voice) user, not increase it. This typically implies more efficient codecs and, in fact, half-rate is already widespread on the local mobile networks for exactly this reason.

On top of this, the whole argument of what is "good enough" bandwidth to carry intelligible speech.

As explained above, Hi-Fi is nice to have but not needed at all and is counter-intuitive to the factors at play, i.e. with a rapidly increasing consumer base and serious constraints on capacity, HD voice is a luxury SA cannot afford at this point in time.
 
Last edited:
Well, one typo on the 70 000 Hz instead of 7000 Hz got people pretty worked up, without looking at the bigger picture.
Maybe I'm missing the bigger picture, but extending the frequency response range doesn't seem all that important even with the harmonic frequencies you pointed out. Granted GSM quality is horrible, but improving quality to beyond what a fixed line does? I'd rather have mobile IP based connectivity that is as stable for carrying voice as GSM is. With a solid IP connection I can have fixed line quality audio, plus all the added functionality and freedom that comes with IP - this is what really matters IMHO.
 
My hi-fi speakers go to 35khz with a ribbon tweeter and I've often wondered why.

I've Googled and apparently people can perceive sounds above 20khz - even if they don't hear them.

I haven't read through this as it's got graphs and kinda scientific looking, but here's something for the technical guys :

http://www.cco.caltech.edu/~boyk/spectra/spectra.htm
 
This post is a mix of verifiable research and personal opinion after 20 years supporting recording studios.

The intelligibility of speech is complex beyond what some posters suggest.

What differentiates the vowel sounds is that the harmonic content differs. A speaker can say each of the vowels at the same pitch - what makes each vowel sound distinct is the relative harmonic energy across several orders of harmonics. For most speakers and listeners a 300 - 3400 Hz bandwidth is adequate for vowel sounds.

Plosives (P sounds) and sibilants (S and T sounds) require respectively lower and higher bandwidth limits to be accurately captured. The correct name for K's and hard C's eludes me for the moment. Plosives and sibilants play a significant role in intelligibility. As bandwidth is reduced the chances of a listener confusing words such as postulate and conjugate increase.

What helps us out very much in phone conversations is that words are not spoken in isolation and we are often able to guess a word according to the context of the sentence even when we are not certain of that word in isolation. This may happen subconsciously.

Dynamic range compression often helps with intelligibility but can be overcooked to the point of harm.

Group delay or phase coherency remains controversial but is increasingly considered a possible factor in intelligibility. This requires better filters or increased bandwidth.

I do know that within a mix situation that restricting spoken voice to the traditional phone bandwidth is harmful to intelligibility. I see no reason why extending the upper limit would not improve phone quality and on the lower side a modern phone that can do (limited) justice to drum and bass can certainly deal with some lower frequencies than 300 Hz.

If the argument is that this is the best compromise I don't really have a beef; if the argument is that dialog requires no more than 300 - 3400 Hz for best intelligibility, well you are wrong!
 
Top
Sign up to the MyBroadband newsletter
X