Building a Vocal Synthesizer from Scratch with Open Audio Models
A conceptual walkthrough: recording consented vocal stems, preparing data, avoiding phase artifacts and getting the result into a plugin your DAW can load.
Sarah Lin
Senior Tech Editor • • 4 min read

The quick take
- 1Consent comes firstRecord your own voice, or work only with singers who have agreed in writing to this specific use.
- 2Data quality winsClean, dry, well labeled stems matter more than clever model settings.
- 3Ship it as a pluginA small wrapper with low latency turns a research experiment into something a producer can actually use.
Few projects are as satisfying as hearing a voice you built come out of your own speakers. Few are also as easy to get wrong. This conceptual walkthrough covers the path from empty folder to a vocal synthesizer that a digital audio workstation, or DAW, can load. It uses no specific product and no code, because the principles outlast any one toolkit. The most important step comes first, and it is not technical.
Step one: consent and ownership
A voice is part of a person. Before you record anything, decide whose voice this will be. The simplest and safest answer is your own. If you want to work with another singer, get a written agreement that names the project, says what the model may be used for, explains whether it may be shared, and describes how the singer can withdraw. This is general guidance, not legal advice, and you should consult a professional for anything commercial.
Never train on a voice you merely admire. Recordings you can download are not recordings you may clone. If a project cannot pass the test of explaining it comfortably to the singer's face, rethink it.
Step two: recording stems
Good input is the biggest lever you have. Aim for a quiet room, a decent microphone and a consistent distance from it. Record dry, with no reverb, delay or heavy compression, because effects baked into the training data will be baked into the voice and cannot be removed later.
Cover range and style. Capture low and high notes, soft and loud singing, and a variety of vowels and consonants. Add spoken passages if you want the voice to speak too. Keep a log of what you recorded and when, and save original files untouched. Think of them as your masters.
Step three: preparing the dataset
Raw recordings become a dataset through patient, unglamorous work.
- Cut the audio into short phrases and remove breaths that were mistakes, along with clicks and noise.
- Normalize levels gently so loud and quiet takes sit in a similar range.
- Label each phrase with text and, if your approach needs it, pitch and timing information.
- Hold back a small portion of phrases that the model never sees, so you can test it fairly afterward.
Resist the urge to add more data of low quality. A smaller set of clean, accurately labeled recordings usually beats a large messy one.
Step four: training at a high level
Open audio models learn a mapping between what you want to say or sing, such as notes and lyrics, and the sound of the voice. You do not need the mathematics to work with them, but you do need the mindset. Training is iterative. You start a run, listen to checkpoints along the way, and stop when the output improves no further or begins to sound overfitted, meaning it merely parrots the training phrases.
Listen on good headphones and on cheap speakers. Check held vowels, fast consonants and the transitions between notes. Keep notes about each run, because three weeks later you will not remember which settings produced the good one.
Common artifacts and how to fight them
- Phase artifacts. A metallic or watery shimmer often comes from the stage that converts predicted features back into a waveform. Try a different vocoder setting, check that your sample rates match throughout the pipeline, and avoid repeated resampling.
- Harsh sibilance. Sharp S sounds can be over-emphasized. A gentle de-esser on the output, or cleaner source recordings, usually helps.
- Pitch wobble. Unsteady notes point to weak pitch labels or too little data in that range.
- Mushy consonants. Timing information that is slightly off will blur the starts of words. Re-check your labels.
Latency and the plugin wrapper
A model that works in a notebook is not yet an instrument. A producer wants to press a key and hear a result without a long wait, and wants the voice to sit in the project like any other track. That requires thinking about latency early. Smaller models respond faster, and processing audio in short blocks reduces delay, though sometimes at a cost in quality. Decide whether your goal is real-time performance or offline rendering, since the trade-offs differ.
Wrapping the model as a plugin is mostly engineering. The wrapper receives notes and lyrics from the DAW, passes them to the model, and returns audio on time. Test it in more than one host, with different buffer sizes, and make sure it saves and restores its state cleanly when a project is reopened.
Sharing, and the ethics of release
Once it works, the harder question arrives: who gets to use it? A voice model is a portable copy of a person's sound. If it is your own voice, you can choose your terms. If it belongs to a collaborator, follow the agreement you made. Consider adding an audible or embedded marker so listeners can tell the output is synthetic, and state clearly in your credits that a voice model was used.
Think about misuse before you release, not after. Share a model only with licenses that forbid impersonation, and keep the training data private. A vocal synthesizer built carefully can be a genuine creative partner. Built carelessly, it becomes a problem for someone else to clean up.
Launch edition note: this is an illustrative story. The studios, platforms, people and events are fictional. See our disclosure protocol.
Sarah Lin
Senior Tech Editor at NewsEntertAI. Launch-edition byline. Spotted an error? Tell us.