How I Learned to Rap in 4 Languages I don’t speak in 1 Night Using the Free Application “Audacity”
Check out this video! Idahosa is back once again for another guest post about using rap to learn a language, and this time he's bringing this absolutely amazing demonstration, where you can hear him rap in eight languages, four of which he doesn't even speak!
Cool as it is, he has broken down the key steps he went through to rap the parts of the song where he was singing in these as-yet-unknown languages, using a really cool and completely free cross-platform application called Audacity.
While I take my hat off to him in terms of his fantastic music, editing and synchronisation skills in the video, his straightforward explanation has me seriously looking into doing some rapping over the next month to help me improve my as-of-yet not fluent and quite choppy Mandarin. I honestly feel like I could personally rap in Japanese (which I don't speak at all) in just a couple of hours after reading this post and his useful audio samples for that language segment of the video!
Have a read for a close-listening and performance exercise. It can sharpen your attention to timing and sound, but a memorised rap does not demonstrate native-like pronunciation or general speaking ability.
In my last guest post, I discussed the benefits of freestyle rap training as a language learning activity. I'm sure many of you read that post and thought: “Sounds cool…but I'll sure as hell never be able to do that.” If that was your mentality, you'll probably think the same way after watching my “Flow Anthem” video (above). That's why I am writing this current post – I aim to prove to you that nothing I did in the above video is beyond your capabilities.
A note from the Fluent in 3 Months team before we get started: You can chat away with a native speaker for at least 15 minutes with the “Fluent in 3 Months” method. All it takes is 90 days. Tap this link to find out more.
The Question
As I explain in this video about The Mimic Method Approach and Technique, your primary goal as a language learner should be to master the sound patterns, or “Flow,” of your target language. With the “Flow” down, you can effortlessly mimic native speech sounds and attach meanings to them as you accumulate more target-language experiences.
Foreign-language mimicry is challenging. Connected speech moves quickly, and familiar sound categories can shape what we think we hear in a new language. Focused listening can help, but it does not remove the need for reliable text and feedback from a proficient speaker.
So assuming that your goal is oral-fluency, the question you have to ask yourself is this:
“How can I learn to hear and speak foreign speech without having to think about it?”
The Traditional Answer: Leximania
No it's not a real word (at least not yet), I just coined it now to make the point that I'm about to make. “Lexi-” means “of or pertaining to words,” and “-mania” means “obsession.” So “Leximania” refers to “an obsession with words.”
Language learners often approach unfamiliar speech through written words.
Everyone approaches the foreign speech problem through words. It seems like a good idea. Normal speech is too fast, so why not break it down word for word and learn each word individually? But as I explained in my post on The Flow of Fluency, words are unreliable language learning tools. Depending on context, a word’s pronunciation can change with speech rate, stress and neighbouring sounds.
This is why so many language-learners complain about not being able to recognize words in normal, connected speech despite having a substantial vocabulary knowledge. What these learners fail to appreciate is that knowing what a word means is not the same thing as knowing how to use it (see wikipedia articles on the difference between Declarative Knowledge and Procedural Knowledge).
Leximania in language-learning is actually just a byproduct of our society's pandemic what-a-mania. When we encounter something unknown, we must to figure out what it is. So for unknown languages, we turn to textbooks that explain all the whats of the language. But language acquisition isn't a “what” activity; it's a “how” activity. Indeed, nobody learns what to speak English; they learn how to speak English.
Similarly, we don't learn what to ride a bike but rather how to ride a bike, and no one learned how to ride a bike from reading a book. The only way to learn is to hop on that baby and start pedalin'.
The Mimic Method Answer: Phonomania
I'll admit to being very what-a-manic about certain things, but when it comes to learning languages, I'm a die-hard Phonomaniac (“Phono” meaning “sound). In fact, I'll go ahead and distance myself even further from the Leximaniacs by spelling the word more phonetically from here on out — “Fonomeniak”.
We Fonomeniaks care little about words and grammar. Our only goal is to master the target-language's sound system, or “Flow”. We do this by looking closely at the sounds of natural speech (not the discombobulated word-for-word speech that Leximaniacs love). Since the goal is to NOT have to think, we practice training our mouths and ears to hear and recreate these sounds automatically.
You could learn the Flow bit by bit starting with simple phrases, but there's really no reason to dilly-dally with baby talk and Dr. Seuss poems. If you want to master the Flow as quickly as possible, you need to dive right in to the the most phonetically complex form of speech – Rap.
Think about that crazy kid on your block growing up who was already popping wheelies within a week of getting his first bike. He wasn't a bicycle savant or anything, he just didn't care about busting his knees and elbows. His willingness to learn the most difficult skills right from the outset turbo-accelerated his learning curve. So by the time you were finally getting your training wheels taken off, he was already riding with no hands and waving at that girl across the street you always had a crush on. That's why I am a strong advocate of learning to rap in your target-language, whether you've been studying for years or just getting started.
Rapping in a language you do not speak can be demanding, but a short passage can be divided into manageable listening and rehearsal tasks. As the Idahosa quartet explains in “The Flow Anthem”:
“…take the sound, break it down – rhythmic, phonetic. You learn the syllables separate, and then connect them together; what you get is…”
Rhythmic Phonetic Training With Audacity
Audacity is a free and open-source audio editor. It is also the Fonemaniaks ultimate language-learning tool. While the Lexomaniac uses the written word to examine the imaginary components of speech (the words), the Fonomeniak uses Audacity to examine the real components of speech (the sounds).
In the sections below, I’ll show you how I used Audacity to isolate, slow and loop short passages. With these techniques, I learnt to perform the foreign-language sections in the second verse of “The Flow Anthem” in two hours. That was a memorisation and performance result, not proof of a near-perfect accent or understanding.
The Basics
You can download Audacity for free here and install it on your computer's hard drive. When you open Audacity, you will see a blank grey workspace and a tool bar at the top.
Use File –> Import –> Audio to choose an audio file you are entitled to use. My Japanese example came from RIP SLYME’s “Nettaiya” (熱帯夜). The old link and spelling were wrong.
As you can see in the screenshot below, the audio data is displayed visually as sound wave, with the horizontal axis representing time, and the vertical axis representing amplitude (loudness). It might look very technical at first, but after some fooling around you'll get used to the interface and find editing the audio as intuitive as editing text on a word processor.

In fact, just like with a word processor, you will rely mostly on the “selector” tool. In the top toolbar, to the right of the red circle “Record” button, you will see six toolbar buttons. The top left button is your selector tool. Select it.

When you click and select a moment of time on a track and press play (space bar), it will playback the audio starting from that select point in time. If you press play again, it will stop the music and go back to where it started.
Isolating your track
Once you decide on which part of the song you want to learn, the next step is to isolate it from the rest of the track so that you can focus on it exclusively. To do this, click your mouse somewhere near the start of your song, then press the zoom icon several times until you have a real close up view of the sound waves.

A waveform shows amplitude over time. Its peaks and quiet points can help you navigate a recording, but they do not map reliably onto syllables, morae or word boundaries. Find edit points by listening repeatedly and checking that you have not cut off a consonant, vowel or transition.
For the Japanese passage in “The Flow Anthem”, I placed the cursor near the first sound, listened several times and adjusted the selection by ear. I then zoomed out and removed the material before that point.
Press play to check whether the excerpt begins cleanly at the intended sound.
(If this isn't playing, click “download” to get the audio. Click X at top-right of embedded audio after listening once to hear it again. Those reading this via RSS or email click through to the site to hear it)
Slowing it Down
Now that we have selected the excerpt, the next task is close listening. A moderate tempo reduction can make a fast passage easier to inspect.
Usually when you slow down a track, you lower the pitch as well since you're physically stretching out the sound wave. That's how you get that clichéd “slow motion voice” effect in movies, when the actor dives to catch a falling pie or something while screaming “Nooooooo.”
Fortunately, Audacity has a special tool for slowing down the audio without altering the pitch. Highlight the whole track (double click or do select all shortcut- ctrl “A”). Then in the top menu bar, select Effect –> Pitch and Tempo –> Change Tempo. This effect allows you to slow down or speed up the audio without effecting the pitch (The “Change Speed” effect DOES alter the pitch).
Depending on the speed of the song and my familiarity with the Flow, I typically reduce the tempo anywhere from 15-45%. Be conservative with this tool, because if you reduce the tempo too much you'll start to distort the speech sounds beyond recognition. The audio below has been reduced 35%.
Identifying the Speech Sounds: Japanese example
Write temporary listening notes for short sections, but do not treat them as a transcript. In Japanese, rhythm is commonly analysed in morae, and neither mora nor syllable boundaries can be read directly from waveform peaks. Loop a short selection, write what you think you hear, then check it against an authorised lyric and a proficient speaker.
My original scratch notation did not pass that check. The official song title and licensed Japanese lyric show that I collapsed long vowels, confused several consonants and imposed unreliable boundaries. Personal spelling can help you remember a first impression, but it should be corrected before you rehearse it.
The original learner-created pseudo-transcript has been removed because it did not reliably represent the Japanese sounds or word boundaries.
Play the whole passage, compare your notes with an authorised lyric, and ask a proficient speaker to check anything you cannot resolve. Keep the corrected version as a reference while you practise.
The notes are a scaffold, not the final exercise. Shift your attention back to the recording once you know what the passage contains, but return to the checked text whenever your memory or pronunciation drifts.
Construction
Next, practise the passage in short rhythmic chunks. Choose boundaries by listening and by checking the lyric, not from the shape of the waveform. Start with a phrase short enough to repeat accurately, then combine phrases gradually.
For each group, listen several times and repeat it to a steady beat. Stop and recheck the source if repetition is reinforcing a doubtful sound.
The examples here originally repeated the unreliable pseudo-transcript in nine chunks. Those spellings have been removed. Use checked lyrics and your own licensed recording to define the practice chunks.
Now that I'm comfy with the bite-size chunks, I group them together into the next level of rhythmic grouping and repeat the same steps.
Combine two checked chunks, compare the result with the original recording, and only then add the next one.
Longer sound sequences place more demands on working memory. If accuracy falls, return to shorter chunks instead of pushing through a fixed syllable count.
It is often easier to repeat a short number than a long one after a single hearing. Chunking applies the same practical idea to a musical phrase.
So for this phase, we have to bust out the more hardcore Audacity tricks.
Memorization:
For this task, we want to create an audio looped file for the song. I've already cut the first part of the song so that my selection starts right on that first syllable. Now I want to cut off the tail end of the audio so that I can copy and paste the tracks next to each other. Here's what I get:
It's important to keep a steady beat throughout the whole thing. As a musically-trained individual, I have a lot of experience thinking about music theoretically and thus do not have a hard time identifying the exact start and end points of a track for it to loop on beat, but for those of you who are not musically trained, you can achieve the same effect with trial and error (hint: ctrl z is the “undo” shortcut).
You'll also notice the use of the “Fade Out” effect (found under in the effects menu again). I've found this aids the memory process by delineating a clear starting point of the lyric and musical beat.
I make the loop file last about 1 minute and sing along with it over and over again until it stops. THEN I try to recall it from memory without the aid of the looped audio. This step is crucial, because the actual recall process is what does most of the work of burying this new info deep into my long-term memory. Once again, I sing the full lyric out loud to myself 30 times in a row to make sure I have it.
I then repeat the process with the next two groupings.
Then for the entire song lyric.
Finally. Back to normal speed (you can gradually build the speed too).
After one 30-minute session, I could perform the short passage from memory. That result says nothing by itself about accurate Japanese pronunciation, comprehension or long-term retention. It shows that looping and chunking helped me memorise this particular musical sequence.
Limitations
As mentioned before, our brains rely heavily on the existing phonetic architecture when processing foreign sounds. This means there is still a good chance that you will hear or create certain songs incorrectly, even if you listen to it several times. This is especially true for foreign language sounds that you have never heard before.
Japanese was unfamiliar to me, and my notes missed distinctions that matter. Japanese rhythm is commonly described in morae; long vowels, the moraic nasal and the small tsu need attention that an English-style syllable count can obscure. Russian and Swahili also included sounds that were new to me. My phonetics and linguistics training did not make an unaudited performance automatically accurate.
Rethinking teaching models
Using Audacity to break down a song can support close listening, but editing time is not language practice. A waveform and repeated listening cannot certify your transcription or pronunciation, so use an authorised lyric and feedback from a proficient speaker.
This is where a trained teacher can be very useful. For my Mimic Method Students, I create the audio materials for learning the songs, they use the materials and learn the songs at their own pace, then they submit recordings of themselves singing the songs. I then give them precise feedback on each sound they mispronounce. Now that they have the main part of the song memorized and don't have to think about it, fixing the few errors here and there is easy.
But as I always tell my students: I do NOT teach language. You don't learn language in class, you learn language in the real world, listening to and mimicking real native sounds. I'm just a sound consultant who can guide you along the path to Flow and Fluency. I can't teach you how to speak Spanish/Portuguese/or Chinese anymore than I can teach you how to ride a bike. Sure, I can hold the handlebars for you while you try to get your balance, but ultimately, it's up to you to just do it.
Social