Chunk-based learning

    What Are Chunks, and Why Native Speakers Don't Translate

    Native speakers aren't fluent because they think faster. They're fluent because a large share of what they say never gets assembled at all — it comes out of memory in one piece. Linguistics has a name for those pieces: chunks.

    8 min read
    Chunky fitting three puzzle pieces together to spell “what are chunks”

    1. What a chunk is: bigger than a word, smaller than a sentence

    Open almost any English coursebook and you'll find two units of learning: a vocabulary list and a set of grammar rules. Behind that pairing sits an assumption nobody states out loud — that speaking is computation. Think of a meaning, select the words, apply the rules, produce the sentence.

    Linguists have been pointing at the same problem with that picture since the 1980s: when native speakers actually talk, the units they work with are much larger than words. A great deal of the time they aren't computing at all. They're retrieving.

    A chunk is one of those retrievable strings — stored whole in the mental lexicon, pulled out without being rebuilt word by word. It might be three words or seven; it might be a complete sentence or a frame with a gap in it. Length isn't the test. Being remembered as one thing is the test.

    Word by word: 8 units to assembleCouldyougivemeahandwiththis?In chunks: 3 units, pulled out wholeCould yougive me a handwith this?
    One sentence, two ways of storing it. Word by word means getting eight units right in the right order; in chunks it's three pieces to line up.

    Take Could you give me a hand with this?. As words, that's eight units, each of which has to be chosen, ordered and inflected correctly — miss one and the sentence stalls. What a native speaker stores is three pieces: Could you, give me a hand, with this. The middle one matters most: give me a hand has nothing to do with hands, so there is no way to derive it. It can only be learned whole.

    2. How much of native speech isn't built on the spot?

    This isn't an impression — there are corpus numbers. Erman and Warren (2000) estimated that prefabricated expressions account for around 58.6% of spoken English and 52.3% of written English. More than half of what a native speaker says was not assembled in the moment.

    Spoken English58.6%Written English52.3%Share made of prefabricated chunks (Erman & Warren, 2000)
    Source: Erman, B. & Warren, B. (2000). The idiom principle and the open choice principle. Text, 20(1), 29–62.

    The first systematic argument for this came from Pawley and Syder (1983), who posed two puzzles that are still quoted today. Why do native speakers reliably pick the one natural phrasing out of the many that would be perfectly grammatical (nativelike selection)? And how do they produce long stretches of connected speech with so few pauses (nativelike fluency)?

    Fluency doesn't come from computing grammar faster. It comes from having a large stock of ready-made, memorised sentence units to draw on.
    The central claim of Pawley & Syder (1983), Two puzzles for linguistic theory

    The corpus linguist John Sinclair (1991) turned this into two coexisting principles: the idiom principle and the open-choice principle. Open choice is the model we all know — combine words freely according to the rules. The idiom principle is retrieval of semi-fixed phrases already stored. Sinclair's argument was that in real language use the idiom principle is the default, and open choice is the fallback for when no ready-made phrase will do.

    3. Why more vocabulary doesn't close the gap

    A lot of learners share one experience: solid reading scores, a few thousand words banked, and still a stall the moment they have to speak. That isn't a discipline problem. It's a structural bottleneck caused by learning at the wrong unit size.

    Translating as you speakThink it in L1Translate the wordsApply grammarSay itPulling a chunkRecognise the situationSay it
    Two routes to the same sentence. Every extra step is another place the sentence can stall.

    Working at word level means running the whole pipeline for every sentence: reach for the meaning in your first language, translate the words, apply the grammar, speak. All of that has to finish inside the two or three seconds a conversation gives you, and any step going wrong sends you back to the start. Working at chunk level lets you skip most of the middle.

    There's processing evidence too. Conklin and Schmitt (2008) compared formulaic sequences against matched non-formulaic strings and found that both native and non-native readers processed the formulaic ones significantly faster. The advantage isn't only that chunks come to mind — they're cheaper to handle once they do.

    On the teaching side, Boers et al. (2006) ran one group through a course with an explicit focus on formulaic sequences and another through conventional instruction; blind raters judged the formulaic group noticeably more fluent in speech. Worth noting: that study had 32 participants, which is small enough that it shouldn't be treated as settled. But its direction agrees with the corpus and processing evidence above.

    4. What chunks look like: four common types

    Chunks aren't one single thing. Nattinger and DeCarrico (1992) and Wray (2002) proposed various taxonomies; in practice four categories cover most of what a learner needs:

    TypeExamplesWhat makes it a chunk
    Collocationsmake a decision, heavy rainWords that habitually co-occur — swapping in a synonym sounds wrong (do a decision is simply not English)
    Idiomsgive me a hand, call it a dayThe meaning can't be derived from the parts, so the whole thing has to be stored
    Sentence framesI was wondering if ..., Would you mind ...ing?Fixed opening, free ending — one frame covers dozens of situations
    Social formulasNice to meet you, Sorry to keep you waitingTied to a situation and used almost without variation
    What they share: all four are strings where storing the whole is cheaper than rebuilding it.

    Sentence frames give the fastest return for speaking. Learn I was wondering if you could ... once and you have the opening for dozens of different requests — you've acquired a whole family of sentences in a single item.

    5. How to actually practise chunks

    The method is usually traced to Michael Lewis's Lexical Approach (1993) and its most-quoted line: language is "grammaticalised lexis, not lexicalised grammar". In daily practice it comes down to four steps:

    1. 1Choose chunks by situation, not from a word list. Start from the settings you'll actually be in — ordering food, an interview, small talk with colleagues. Chunks from one situation reinforce each other, which beats memorising scattered items.
    2. 2Say the whole thing out loud; don't break it up. Half a chunk's value is its rhythm. give me a hand is one rhythmic unit; pronounced as four separate words it loses the very cue that makes it retrievable as a whole.
    3. 3Use shadowing to get the intonation. Hear a line, say it back immediately. What you're copying isn't only the sounds — it's the stress placement and the linking. This is the step that moves a chunk from *recognised* to *sayable*.
    4. 4Review on a spacing schedule. Chunk memory decays like any other. Coming back after a day and again after three beats one thirty-minute session.

    Steps two through four are what ChunksEnglish is built to do: native audio on every chunk, shadowing with recording playback, and spaced review. The method doesn't depend on any app, though — a notebook and YouTube will get you there, just with more friction.

    6. Three common misconceptions

    "Chunks are just sentence patterns"

    Traditional pattern drills are a structure plus a list of substitutable words, and the point is the grammatical structure. Chunks are about habitual co-occurrence — which words native speakers actually put together. heavy rain is grammatically unremarkable, but it's a chunk, because nobody says strong rain.

    "Learning chunks will make me sound robotic"

    The opposite. Because the fixed part is automatic, your attention is free for the small part that carries what you actually mean. Cognitive resources are finite; the effort you don't spend assembling sentences is effort you can spend on content.

    "This is only for beginners"

    The chunk gap is more visible at higher levels, not less. Beginners are limited by vocabulary. Intermediate and advanced learners are usually the ones who can say it but still don't sound native — which is exactly Pawley and Syder's nativelike selection problem, and chunks are the only thing that fixes it.

    7. Where to start

    The most effective opening move is to pick one situation you'll genuinely be in this month and take its chunks as a set, rather than collecting them at random. Each list below gives every chunk with its meaning and three examples in context:

    Take one situation, drill it until the lines arrive without thinking, then move on. That finishes far better than opening four situations and learning ten items in each.

    Frequently asked questions

    What's the difference between a chunk and an idiom?

    An idiom is one kind of chunk, but chunks cover far more ground. make a decision isn't an idiom — its meaning is entirely transparent — yet it's still a chunk, because that's the pairing native speakers use and do a decision isn't English. The test isn't whether the meaning is derivable. It's whether speakers treat the string as one unit.

    How many chunks a day is enough?

    Better question: have you finished a situation? Five a day, one situation per week, beats twenty a day spread half-finished across four situations — chunks only carry a conversation when you have the set. The real threshold is that a line arrives without you reaching for it, not a count.

    Should I memorise the translation too?

    Early on, yes — but the translation is scaffolding for understanding, not the memory target. The link you're building is situation → English chunk, not first language → English. So cue yourself with the situation (you walk in, the server approaches) rather than reading a translation and converting it. The second one drills exactly the habit you're trying to lose.

    I'm already at an intermediate level. Do I still need chunks?

    More than before, usually. Beginners are held back by vocabulary size. Intermediate and advanced learners are typically held back by "grammatically fine, but no native speaker would phrase it that way" — which is precisely Pawley and Syder's nativelike selection problem, and only accumulated chunks close it.

    Does this help with IELTS or TOEFL?

    Directly for speaking and writing, indirectly for reading and listening. The IELTS speaking band descriptor for Lexical Resource explicitly rewards natural collocation and idiomatic usage — exactly what chunks are. If your goal is a short-term reading score, though, straight vocabulary work is faster.

    Can I practise chunks without an app?

    Yes. The method doesn't depend on tools: take real native material (shows, podcasts, interviews), write down the recurring lines whole, say them whole, come back to them a few days later. What an app saves you is finding material, cutting it into lines, and scheduling review — not the method itself.

    References

    1. 1.Pawley, A. & Syder, F. H. (1983). Two puzzles for linguistic theory: nativelike selection and nativelike fluency. In Richards & Schmidt (eds.), Language and Communication. Longman.
    2. 2.Erman, B. & Warren, B. (2000). The idiom principle and the open choice principle. Text — Interdisciplinary Journal for the Study of Discourse, 20(1), 29–62.
    3. 3.Sinclair, J. (1991). Corpus, Concordance, Collocation. Oxford University Press.
    4. 4.Lewis, M. (1993). The Lexical Approach: The State of ELT and a Way Forward. Language Teaching Publications.
    5. 5.Wray, A. (2002). Formulaic Language and the Lexicon. Cambridge University Press.
    6. 6.Conklin, K. & Schmitt, N. (2008). Formulaic sequences: are they processed more quickly than nonformulaic language by native and nonnative speakers? Applied Linguistics, 29(1), 72–89.
    7. 7.Boers, F., Eyckmans, J., Kappel, J., Stengers, H. & Demecheleer, M. (2006). Formulaic sequences and perceived oral proficiency: putting a lexical approach to the test. Language Teaching Research, 10(3), 245–261.
    8. 8.Nattinger, J. R. & DeCarrico, J. S. (1992). Lexical Phrases and Language Teaching. Oxford University Press.
    9. 9.Miller, G. A. (1956). The magical number seven, plus or minus two: some limits on our capacity for processing information. Psychological Review, 63(2), 81–97.

    Practise in chunks, two minutes a day

    ChunksEnglish organises chunks across 20+ everyday situations, each with native audio, shadowing and recording playback. Core features are free.