On this path · Speaking 4. Connected Speech
  1. Rhythm and Sentence Stress
  2. Word Stress and Clarity
  3. Intonation: Meaning in the Melody
  4. Connected Speech
  5. Pronunciation for Clarity
  6. Fillers and Thinking Aloud
  7. Conversation Flow and Turn-Taking

Connected Speech

Link, reduce, and elide sounds so speech flows like a native's.

Learning outcomes

You learned, in rhythm and sentence stress, that English squeezes the unstressed words to keep a steady beat. This page follows that squeeze all the way down to the level of individual sounds. When words get crushed together by the rhythm, their sounds do not stay neatly separate: they blend, weaken, and drop out. That is connected speech, and it is the single biggest reason fast native English is hard to follow, even when you know every word on the page.

After studying this page, you can:

  • Explain why words do not stay separate in natural speech, and name the main processes that fuse them: weak forms, linking, elision, and assimilation.
  • Hear and produce weak forms, where common function words collapse around the schwa, and recognize the schwa as the engine of the whole system.
  • Identify linking, elision, and assimilation in real workplace and everyday phrases, and respell what a sentence actually sounds like.
  • Reframe connected speech as a listening problem first, so that fast native speech becomes something you can segment rather than a wall of noise.
  • Judge when to reduce and when to slow down, and avoid the three traps: robotic word-by-word delivery, mishearing blended speech, and never reducing function words at all.

Before we dive in

Here is a moment every learner knows. A colleague leans over and says something that sounds like “jeet yet?” You freeze. You know thousands of English words and not one of them is “jeet.” Then it lands: “Did you eat yet?” Every word was ordinary. The problem was never vocabulary. The problem was that the words did not arrive as separate words. They arrived fused, weakened, and partly dropped, the way English actually comes out of a mouth at speed.

This is the gap that connected speech fills. Textbooks teach words in their full, careful, dictionary form, one at a time, each sound pronounced. Real speakers almost never talk that way. They run words together, lean hard on the few that carry meaning, and let the rest collapse. The careful forms are real and useful, but they are the slow-motion version. Connected speech is the same language at normal speed, and the changes are not sloppiness or bad habits. They are systematic, they follow rules, and native speakers produce them without thinking.

A note on how we will write sounds. We will not use phonetic symbols. Instead we respell words the way they actually sound, in plain letters, and put the STRESSED word in capitals where it helps. So “want to” gets respelled as “wanna,” “did you” as “didja,” and “next door” as “nexdoor.” Read the respellings out loud; that is the whole point. The vocabulary here, weak form, schwa, linking, elision, assimilation, is just a set of labels for things your ear has heard a thousand times without a name.

The five things that happen at speed

Connected speech is not one effect but a small family of them, and it helps to keep them apart because each does something different to the sound. The table below is your map. It shows the full written phrase, what it actually sounds like when spoken at speed, and which process did the damage. Read the middle column aloud and you will recognize most of these instantly.

The same phrase in its written form, in its connected-speech sound, and the process that reshaped it.

Written formHow it sounds (respelled)Process
fish and chipsfish-n-chipsWeak form
a cup of teaa cup-uhv teaWeak form
I have to goI haff-tuh goWeak form
an applea-nappleLinking
pick it uppi-ki-tupLinking
law and orderlaw-r-n-orderLinking (intrusive r)
next doornexdoorElision
I don't knowI dunnoElision
comfortablecomftableElision
ten bikestem bikesAssimilation
did youdidjaAssimilation
won't youwonchuAssimilation

Notice that the written and spoken columns can look almost unrelated. “Next door” loses a whole sound; “did you” grows a “j” that is in neither word; “ten bikes” turns its “n” into an “m.” A learner who is listening for the written forms will never find them, because the written forms are not what gets said. The rest of this page takes each process in turn, so you can both recognize it when you hear it and use a little of it to sound less robotic.

1. Weak forms: the function words that almost disappear

Start with the most common process of all, because it touches nearly every sentence. English has two kinds of words. Content words carry the meaning: nouns, main verbs, adjectives. Function words are the grammatical glue: “and,” “to,” “of,” “for,” “can,” “you,” “have,” “a,” “the.” In connected speech, content words stay strong and function words go weak. A weak form is the squashed, unstressed version of a function word, and almost always its vowel collapses into a tiny, neutral “uh” sound.

Listen to what happens to these common ones. “And” becomes “n,” so “fish and chips” is “fish-n-chips.” “To” becomes “tuh,” so “I want to see it” is “I wanna see it.” “Of” becomes “uhv,” so “a cup of tea” is “a cup-uhv tea.” “For” becomes “fuh,” so “this is for you” is “this is fuh you.” “Can” becomes “kn,” so “I can help” is “I kn help.” “You” softens to “yuh,” and “have” in “have to” weakens to “uhv,” giving the “haff-tuh” you heard above. None of these is careless. The strong, full forms (“AND,” “TOO,” “OFF”) are reserved for emphasis or for saying the word alone; in the flow of a sentence, the weak form is the correct, native default.

The second version is not faster slang; it is what a native speaker says by default. The strong words (“SEND,” “FREE”) carry the message, and the function words shrink out of the way. If you say every word with full strength, as in the first line, you do not sound clearer. You sound mechanical, and oddly enough you are harder to follow, because your listener cannot tell from the rhythm which words actually matter.

2. The schwa: the engine behind all of it

You may have noticed the same little “uh” sound appearing again and again above: in “tuh,” “uhv,” “fuh,” “a cup-uhv.” That sound has a name, the schwa, and it is the single most important vowel in spoken English. The schwa is the relaxed, neutral vowel you make when your mouth does almost nothing, the sound in the middle of “sofa” or at the start of “ago.” It is the destination that every weakened vowel falls toward.

This is the deep reason weak forms exist. Reducing a function word is not random; it is the vowel sliding to the schwa because the word is unstressed. “To” has a clear “oo” when you say it alone, but unstressed it relaxes to “tuh.” “For” has a full vowel alone, but unstressed it relaxes to “fuh.” The same thing happens inside longer content words: in “comfortable,” the middle vowels relax so far they almost vanish, which is why it comes out “comftable.” Once you hear the schwa as the place unstressed vowels go to rest, weak forms stop looking like a list to memorize and start looking like one simple rule: no stress, so the vowel goes neutral. This is the sound-level echo of the rhythm you met in rhythm and sentence stress, where the unstressed syllables get compressed to keep the beat steady.

One vowel does most of the reducing

Almost every reduction in English ends at the same place: the schwa, that lazy neutral “uh.” Weak forms are just function words whose vowel has fallen to the schwa because nothing is stressing them. If you train your ear to expect a schwa wherever a word is unstressed, you stop hearing fast speech as a blur of mystery vowels and start hearing a predictable, relaxed default.

3. Linking: where one word ends and the next begins

So far we have squeezed words. Now we join them. Linking is what happens when a word ending in a consonant runs straight into a word starting with a vowel: the boundary between them disappears, and the final consonant attaches to the front of the next word. We do this because it is easier to keep the voice flowing than to stop and restart for every word.

The classic case is “an apple.” Said naturally, the “n” of “an” jumps onto “apple,” so it sounds like “a-napple.” The same thing makes “pick it up” come out as “pi-ki-tup”: the “k” links into “it,” the “t” links into “up.” This is why a phrase you know perfectly well in writing can be unrecognizable by ear, the gaps you expect between the words simply are not there. English does not leave little silences between words; it glues them at the seams.

British English adds a twist worth knowing, because you will hear it constantly. When a word ends in an “r” sound that is written but not normally pronounced, and the next word starts with a vowel, that “r” comes back to do the linking job. “Far away” becomes “far-r-away.” This is called linking r. Speakers then extend the same trick even where there is no “r” in the spelling at all, inserting one purely to bridge two vowels: “law and order” becomes “law-r-n-order,” and “I saw it” can become “I saw-r-it.” That invented “r” is called intrusive r. You do not need to produce it, but you should recognize it, or “law-r-and-order” will sound like a word you have never met.

4. Elision: the sound that drops out

Linking joins sounds; elision deletes one. Elision is the disappearance of a sound, most often a “t” or a “d,” when it gets caught between two other consonants and becomes too much effort to say. The mouth takes the shortcut and simply leaves it out.

Hear it in “next door.” The “t” sits between the “x” sound and the “d,” and in normal speech it vanishes: “nexdoor.” The same fate hits the “d” in “I don’t know,” which collapses to the famous “I dunno,” and the “t” in “last time,” which becomes “lasstime.” Elision also happens inside single words: “comfortable” loses its middle syllable to become “comftable,” and “vegetable” becomes “vegtable.” The dropped sound is not mispronunciation; trying to force every “t” and “d” back in is what would sound unnatural and over-careful. For listening, elision is the cruel one, because you are waiting for a sound that never comes, so the word never matches what you expected.

Elision in the wild

Say these at full speed and listen for the missing sound. “I must go” becomes “I muss go.” “Old man” becomes “ol man.” “Friendship” becomes “frenship.” “Sandwich” becomes “samwich.” In every case a “t,” “d,” or “n” between consonants has quietly dropped out. Once you know to expect the gap, these stop being mysteries and become ordinary.

5. Assimilation: the sound that changes its neighbor

The last process is the sneakiest, because nothing is dropped or joined: a sound simply changes to be more like the sound next to it. Assimilation is a sound shifting to match its neighbor, so the mouth can glide from one to the next without resetting. It usually works backward, with a sound changing to anticipate the one that follows.

Take “ten bikes.” To say the “b” of “bikes,” your lips must come together, so the “n” of “ten,” anticipating those closed lips, turns into an “m”: “tem bikes.” The same happens in “ten people” (“tem people”) and “good boy” (“goob boy”). A second, very common pattern fuses a “d,” “t,” or “s” with a following “y” sound into a “j” or “ch.” This is the source of “did you” becoming “didja,” “would you” becoming “wouldja,” “won’t you” becoming “wonchu,” and “miss you” becoming “mishyou.” Once you know this pattern, a whole family of fast phrases that used to sound like single mystery words (“whatcha doing,” “gotcha”) snap into focus as ordinary “what are you” and “got you.”

A speaker says something that sounds like 'I'll meet you ou-chere in tem minutes.' Which two connected-speech processes are at work in 'ou-chere' (out here) and 'tem' (ten)?

Contractions and the casual forms

There is one more layer, and you already half-know it. Contractions (“I’m,” “don’t,” “we’ll,” “she’s”) are connected speech that has become so standard it earned its own spelling. They are weak forms and elisions frozen into the writing system, which is why they are acceptable in writing where raw respellings are not. Beyond them sits a set of very common reductions that have spellings used to imitate speech: “going to” becomes “gonna,” “want to” becomes “wanna,” “got to” becomes “gotta,” “give me” becomes “gimme,” “let me” becomes “lemme,” and “kind of” becomes “kinda.”

These casual forms are perfectly natural to say. Every native speaker produces “gonna” and “wanna” without a second thought, in meetings and in casual talk alike. The catch is purely about writing them down. You say them freely, but you spell them out as the full forms in any professional email, document, or comment, writing “going to” and “want to,” and reserve “gonna” and “wanna” for quoting speech or for deliberately casual chat. For the full rule on when these spellings are allowed on the page and when they read as careless, see conversational shortcuts and vagueness.

A whole sentence, written and respelled

Each process is clear on its own. The real test is hearing several of them stack up in a single ordinary sentence, which is exactly what happens in real speech. Below is one plain workplace sentence written normally, then respelled the way it actually leaves a native speaker’s mouth, with the stressed words in capitals.

Look at how little of the spoken line matches the written one. A learner expecting “What are you going to do” word by word hears “WHAddaya GONNA do” and finds nothing to hold onto. Yet there is nothing exotic here; this is a neutral, everyday sentence at normal speed. The lesson is not that native speakers are careless. It is that the written form and the spoken form are two different things, and fluency means knowing both.

Mental Model: connected speech is a listening skill first

The natural assumption is that connected speech is about speaking: a set of advanced tricks to make your own English sound smoother. That is true but secondary, and treating it as only a speaking skill leaves the bigger problem unsolved. Most learners can already be understood. What defeats them is understanding others, and that is where connected speech does its real damage.

The better model is to see connected speech as a listening skill first. The reason fast native speech feels impossible is not speed in itself; it is that your ear is searching for the careful, separate, fully pronounced words you were taught, and those words are not in the signal. The native speaker did not say “did you eat yet” as four clean words. They said “jeechet.” If your ear only knows the textbook forms, you cannot segment the stream into words at all, and the whole sentence dissolves into noise. Learning the processes on this page is really learning what to expect, so that “jeechet” resolves into “did you eat yet” the instant you hear it.

You parse what you predict

Listening to fast speech is not passive reception; it is prediction. Your brain segments the stream by matching it against the forms it expects. If you only expect careful citation forms, blended speech has nothing to match and you hear mush. Once you expect weak forms, linking, elision, and assimilation, the same stream becomes parseable, because now you are predicting the shapes that are actually there. Connected speech is a comprehension skill wearing a pronunciation costume.

The speaking payoff then follows naturally, and it is real. A little linking and reduction makes you sound markedly more fluent than careful, word-by-word delivery. The irony is that over-precise speech, every word equally stressed and fully pronounced, is not actually clearer for your listener. It strips out the rhythm that tells them which words matter, so it sounds robotic and is genuinely harder to follow. You do not need to chase every reduction a native uses. Reducing your function words and linking across a few obvious vowel boundaries already moves you most of the way from robotic to natural.

When to reduce less

Connected speech has a clear boundary, and respecting it is part of mastery. The processes here are the default for relaxed, normal-tempo speech among people who share context. They are not a target to maximize. The right amount of reduction depends on the situation, and several situations call for less.

Slow down and pronounce more fully when clarity is genuinely at risk: when you are giving a phone number, a name, or a precise figure; when you are speaking to someone whose English is still developing; over a bad connection; or in a noisy room. Reduce less in very formal, careful, or emphatic speech, a courtroom statement, a clear public announcement, the moment you stress a single crucial word. And when a word boundary actually carries meaning that blending would erase, keep it crisp. The rule underneath all of this is simple: clarity first. Reduction is a feature of comfortable communication, not a goal that outranks being understood. For the wider set of choices about making each sound land when it counts, see pronunciation for clarity.

Common mistakes

The first mistake is the speaking trap: careful, word-by-word delivery with every word fully pronounced and equally stressed. It feels safe and correct, but it sounds robotic, drains the rhythm that signals meaning, and is paradoxically harder to follow than natural speech. The fix is not to rush; it is to let function words go weak and to link across obvious vowel boundaries, so the stressed words can stand out.

The second mistake is the listening trap, and it is the one that holds most learners back: expecting every word to be said in full, and so being unable to segment fast speech. If your ear is hunting for “did,” “you,” “eat,” “yet” as four clean words, you will never catch “jeechet.” The fix is to train the ear on the processes deliberately, so that the blended forms become the forms you predict.

The third mistake is never reducing function words at all, treating “to,” “of,” “and,” and “for” as if they deserved full stress. This is really the first mistake in close-up: it is the habit that keeps “I want to go” from ever becoming “I wanna go,” and it is the difference between sounding like a textbook and sounding like a person. Let the small grammatical words shrink. They are supposed to.

Mastery Questions

Question

Why do words not stay separate in natural speech, and what are the main processes that fuse them?

Answer

English keeps a steady rhythm by leaning on content words and crushing the unstressed words between them, and at the level of individual sounds that crushing reshapes the words. Four main processes do the work. Weak forms: function words like 'to', 'of', 'and', 'for', 'can', 'you' lose their full vowel and reduce around the schwa ('to' sounds like 'tuh', 'of' like 'uhv'). Linking: a final consonant joins a following vowel so the boundary vanishes ('an apple' sounds like 'a-napple'). Elision: a sound, usually a 't' or 'd' between consonants, drops out ('next door' sounds like 'nexdoor'). Assimilation: a sound changes to match its neighbor ('ten bikes' sounds like 'tem bikes', 'did you' like 'didja'). None of these is careless; they are systematic and native-default.

1 / 5

Recommended next

  • Fillers and Thinking Aloud
    Once the learner can blend and reduce, the next layer is managing the flow of speech itself, including fillers and pauses.
  • Pronunciation for Clarity
    Connected speech ends on the boundary of clarity first; pronunciation for clarity teaches making each sound land when reduction is not appropriate.
Sources & evidence7 claims · 4 cited

Weak forms, the schwa, linking (including British linking and intrusive r), elision, and assimilation are each defined and shown on a respelled case; Roach backs the processes, Wells and Carter and McCarthy the framing, and Biber the frequency of reduced function words. No invented figures; phonetics rendered as plain-English respelling.

  • In connected speech, common function words (and, to, of, for, can, you, have) take weak forms in which their vowel reduces to the schwa when unstressed, while content words keep their strong forms.verified
  • The schwa is the neutral unstressed vowel of English and the destination of vowel reduction, so weak forms and the hollowing of unstressed syllables both arise from vowels relaxing toward the schwa.verified
  • Linking joins a word-final consonant to a following word-initial vowel so the boundary disappears, and in non-rhotic British English a linking r reappears between a written but normally silent r and a following vowel, extended to an intrusive r between vowels with no r in the spelling.verified
  • Elision deletes a sound, most often a /t/ or /d/, when it occurs between consonants (next door to nexdoor, I don't know to I dunno), and can also drop a medial syllable within a word (comfortable to comftable).verified
  • Assimilation changes a sound to resemble its neighbor, so the n of ten bikes becomes m before b (tem bikes) and a d or t plus y fuses to an affricate (did you to didja, won't you to wonchu).verified
  • Because listeners segment fast speech by matching it against predicted word forms, learners who expect every word in its careful citation form cannot parse blended native speech, making connected speech a comprehension problem as much as a production one.internal reasoning
  • Reduction is the default for relaxed normal-tempo speech but should be lessened when clarity is at risk (precise figures, developing listeners, noise, formal or emphatic speech), because clarity outranks reduction.stable common knowledge

Cited sources