Originally presented as a live talk on September 30, 2026
Background
This is a speech model right? So we gotta talk about audio. But to do that, I first want to talk about text, which we’re all more familiar with when it comes to machine learning.
Because it turns out that text is a bit of an outlier in the data universe. Unlike audio or image, or even proprioception and action in the world of robotics, text is a discrete abstraction.
“Discrete abstraction”, what do I mean by that. Well let’s break it down. Text is discrete because everything is individual words or characters. There are clear divisions between minimal units. No judgment involved to split up a sentence into words or words into letters.
That is not the case for audio and image data. And I’m going to talk about both, because image data is more intuitive than audio data but has a lot of the same structure. Anyway, when you divide up some audio or an image, you have to choose where to slice within that data. So like a single image for example, if you break it up into pixels you have to pick a pixel size, and some pixels are going to fall right on the border of two things in the image. But a pixel can only be one color, so you end up with these artificial breaks or discontinuities.
Same deal with audio. Within a sample of speech let’s say, what you’re getting is just a waveform over time. Yes, we know that usually the speech is a series of discrete words, but that is our understanding on top of what is objectively going on. And if you want to divide up that waveform into words or sounds - “phonemes” as the linguists call them - you have to pick out the border between two phonemes and split. And that border requires some judgment, because sound is continuous.
And the reason sound and images are continuous is that they are physical, they are part of the real world. Reality is almost always continuous at the macro level. Like it’s hard for the stripes on a zebra to have sharp edges, it’s hard for a hot thing to be near a cold thing with no warm zone intermediate, it’s hard for a sound to instantly start or stop. Things might seem discontinuous to us, but reality has a surprising amount of detail, and if you look closely there is usually some bleed or a curve rather than a jump.
Abstractions don’t have to be like this. In fact it’s much easier when they aren’t. That way you can mix and match them cleanly, like Lego castles rather than sandcastles. In English, we get letters we can arrange into phonemes and words that represent sound. We lose a lot of detail about that sound, but abstractions are often lossy like that. It’s the price you pay for the nice clean features.
Computers deal in discrete abstractions. At base they are electrical circuits in the physical world, yes, but we have worked to introduce abstractions on top of those circuits as early as possible, so that we can understand and program them.
So text is a natural fit for computers with their 1s and 0s. And if you know about the modern transformer-style LLM, you know it spits out text based on a list of all possible tokens and the likelihood of each one of them being next. That doesn’t work if the output has to be continuous.
Now there is ML machinery to generate continuous-seeming outputs. For example, diffusion models - which draw inspiration from the physical process of diffusion - are common for generating images. But most of the focus is on transformers these days, and transformers generally output probability distributions over discrete options.
There are ways to get transformers working with continuous data though. If it’s just input, you actually can avoid the jump to discrete if you bypass the tokenizer and go straight to the embedding space, where meaning really is continuous over the values of the vector in the embedding space.
Another option is to make the jump, to take the continuous data and tokenize it anyway, even though it doesn’t inherently have the nice clean boundaries that text does. We don’t need to get into audio tokenization for this paper, but basically you can boil down a bit of audio into what sound it represents and how to construct that sound. The graphic here gives an idea of how involved speech tokenization is.
Besides faithfully tokenizing audio, another challenge for model makers is two-way communication and fluidity, as opposed to the turn-based flow of a text chat.
The technical term here is “full duplex”. A full-duplex model allows simultaneous communication from both sides, like a phone or a Zoom. By contrast, a half-duplex model requires one side to wait for the other before replying, like a walkie-talkie.
If you have used a voice-enabled model, it was very likely half duplex. Certainly if you took my ML diagnostic back in April, you experienced half duplex, specifically with the gpt-realtime model family. Half duplex is unnatural for people and results in some clunky user experiences. It also requires the model to sense when you start talking and immediately shut itself up, a technique called “voice activity detection” (VAD). VAD is notoriously finicky and frequently leads to either the model suddenly going silent (VAD too sensitive) or completely ignoring your interjection (VAD not sensitive enough).
Full-duplex models, on the other hand, can gracefully handle interjection and overlapping speech because they can understand what the user sounds mean for the conversation. Like if the sound is just a dog barking in the background, or the user throwing in an “uh huh”, a full-duplex model can hear and understand it while it continues talking. The result is a much more natural conversation. However, full duplex is harder and thus relatively rare. It was a big part of the pitch for the Interaction model from Thinking Machines, which had these impressive demos where a user conversed with the model while the model continued working on its code or its spreadsheet or whatever.
The graphic here is from Thinking Machines by the way. It shows the model speaking even when there is audio and video input at the same time. You can see many more demos and possible use cases on the launch page.
Here is a pioneering full–duplex model from 2024 called Moshi, which the authors of today’s paper will take as the base for SteerDuplex.
Now we aren’t going to cover everything here, but a couple details different from the standard LLM will be crucial to understand.
One is the three-track thing going on at the bottom, with the user’s audio input in purple and two model streams in orange. The first orange one is just Moshi’s audio output, what it is saying back to the user - fine. The other one they call the “inner monologue”, just the text that Moshi is about to speak. So if Moshi’s inner monologue puts out a word in step s there, then in step s+1 it will start speaking the word, continuing to produce audio output tokens in these 80 ms-long steps until it has finished the current word and is ready to speak the next word. And for reference, people generally speak 2-3 words per second. That’s 4-6ish steps per word, meaning most inner monologue steps will be empty, waiting for the sound of the current word to finish so it can kick off a new word.
One interesting detail about how the model produces these sounds is at the top: the distinction between semantic and acoustic output tokens. It’s the difference between what sound the model wants to make, like what phoneme, and how it’s going to make it. They mention in the Moshi paper that for every semantic token the model produces, it also produces seven acoustic tokens, meaning it takes a lot more information to construct the sound than to indicate what sound Moshi should make. Again, you have a tidy little abstraction versus a big complicated physical phenomenon.
Also, speech is information-dense, much more so than I think many English speakers realize. Tonal language speakers may have more of an appreciation, but still, speech is so natural to us all that we take for granted just how much it can convey.
Tech companies are well aware though, and as speech understanding improves it’s opening a lot of doors for product plays.
A popular one nowadays is voice interaction with your computer, like Wispr Flow. But that’s more speech to text, also known as automatic speech recognition (ASR), with AI making it significantly more accurate.
Another popular category is wearables, which can’t work at all without some alternative to the keyboard. At Meta Connect recently, Zuck debuted audio-only smart glasses, in addition to a new version of the existing video+audio smart glasses. He also led the Muse section of the event by announcing Muse would soon support voice, and he closed his related Twitter thread with a shot of this device here, Muse Charm. If you think about it, it’s really no surprise that the socials company now named after the metaverse would go hard on speech tech; voice is a highly interpersonal medium, and typing just doesn’t jibe with the controllers or haptics of VR.
So for the version of AI that Iron Man gets with JARVIS, we really need to master speech tech.
The Paper
So for SteerDuplex, the goal is to train and measure a model for steerability, which the authors define as “the ability to reliably shift conversational behavior along attributes such as tone, persona, speaking rate, and voice style in response to user instructions.” Contrast that with instruction following, which is more about conversational outcomes.
To start off, the authors helpfully orient us on the terms that fall under “steerability”, as well as under “duplex”. We will see this capabilities taxonomy crop up again, so it’s worth spending some time here.
Broadly we have three categories: text, audio, and duplex. Text and audio are about adapting to information in the relevant form, although remember that the actual input is always speech. Here we’re distinguishing between information that is most naturally text and information that is most naturally audio. So an instruction to be more sarcastic is easy to convey in text, but a sigh is a better fit for audio. I’m a little unclear on why the text inputs for Voice/Style Editing and Audio-Context Memory are in the Audio column frankly.
More clearly, anything specific to full duplex is over on the right: taking turns in a natural and timely fashion, handling interruptions from the user, adding affirmations while the user is speaking to show the agent is still there, etc.
An overlapping set of authors published an earlier paper, Audio MultiChallenge, that has overlapping skills in it. We’ll see results on AudioMC later, but keep these taxonomy entries in mind: Audio-Context Memory, Self-Coherence Over Audio, Mid-Utterance Corrections, and Dynamic Mid-Conv. Steering. I have marked them with stars for easy reference.
So that’s the benchmark side. Now here’s the model side, which will of course run against the paper’s benchmark and use some excellent Scale data.
The starting point for the SteerDuplex model is Moshi, the full-duplex model we covered in the Background slides. You can see in part A the parallel tracks for the user and the agent, and the split between semantic and acoustic tokens. They also show the agent’s text output, which Moshi called the “inner monologue”.
I’m going to skip over part B, which SteerDuplex also inherits from Moshi, so that we can focus on the data and training.
Parts C and D show how training happened: one stage of SFT, then two stages of RL, which incorporated different rewards.
For SFT, it’s about 8,500 hours of audio across more than 500k examples, along with a bit of pure text data. We don’t get a full accounting, but the audio data is a mix between natural conversations and examples specifically targeting instruction following, steering, duplex interaction, safety, and reasoning. They also note the natural variation in prosody - a fancy linguistic term for the rhythm, tone, stress and so on in speech.
As for sourcing, some of it is licensed from outside, but a lot of it is Scale data. There is some synthetic speech in here too.
Now let’s talk about RL. Sadly, there is no information about how much data they used. However, we do get some details about what they’re testing and looking for in the model’s responses:
Verifiable timing: objective checks for appropriate pause timing (wait instead of barging in), turn timing (respond when the user finishes), backchannel timing, interruption yield and recovery (stop when interrupted, then respond sensibly), noise robustness (don't react to non-user sound), and speech-mirror .
Transcript rubric: judging whether the text of the response fully addresses the user prompt. No audio involved.
Response continuity: this is basically an anti-reward hacking term, to balance out the model getting points on verifiable timing with short throwaway responses like “Sure” or “Okay”. They put the minimum duration for a substantive response at four seconds.
Continuation: keeps the model responding when the user backchannels, e.g. saying “Uh huh” in the middle of a model response, or when there is unrelated noise in the user’s background. Continuation comes in the second RL stage only, with significant weighting, because after the first RL stage the main failure mode was the model yielding the floor too readily.
They also include a few checks on the waveform itself to filter out clips of pure silence or poor audio.
So now that we have our model, we want to measure it on what we trained it for. That’s the impetus for SteerBench, a novel benchmark testing for steerability in full-duplex models.
SteerBench comprises 390 tasks across tone, persona, style/accent, and speed/length. Each audio prompt uses one of a small set of neutral synthetic voices for uniformity. The task comes with a human-authored rubric, with some criteria for the audio and some criteria for the transcript. There are just over 1k criteria, so the average rubric has 2-3 criteria, which is much lower than I would have expected, likely due to the relative simplicity of the tasks compared to something like GDPval.
For the audio criteria, it’s actually kind of a mix between rubrics and side-by-side evals; there is a text description, like “uses sarcasm”, but there is also a reference audio clip exhibiting the criterion in question. It was an empirical decision to include the reference clips; judge models agreed with human ratings almost 15% more often when they had a reference clip accompanying the relevant audio criterion.
As for the judging, they have three separate metrics:
Audio Pass Rate (APR): the share of tasks where the model satisfies every audio criterion
Sample APR: the share of tasks where the model satisfies every audio and text criterion
Pass Rate: the share of criteria the model satisfied across all rubrics
Here’s where we start to see them in action. On the left we compare two full-duplex models to SteerDuplex after SFT but before RL. As the name suggests, SteerDuplex is more steerable, and the gains come from SFT. (We’ll see the benefits of RL later.)
On the right we have the same model on the other two metrics, although they call Sample APR “All constraints” instead. Not only is speed/length the hardest to get right individually, it’s also significantly harder to get entirely. There also seems to be plenty of room to improve on the benchmark, which as a benchmark author must be a relief to see.
Still on SFT results, we get a look at improvement on a few other benchmarks.
For AudioMC, first let’s recap what it measures:
Inference Memory (IM): the ability to use information from earlier in the conversation to tailor the current response
Instruction Retention (IR): whether or not the model keeps following standing orders
Self Coherence (SC): sticking to earlier claims from the model itself, rather than flip-flopping due to user input
Voice Editing (VE): keeping user context and instructions straight after the user makes changes
For smaller models in particular, which most audio models are due to latency concerns, all four present challenges. But it seems our audio data really helps. These are hardly SOTA results, but the point is relative improvement over the Moshi baseline.
For Full-Duplex-Bench (FDB) v2, the setup is GPT Realtime as the examiner and the model in question as the evaluatee. There are a few different categories of scenario, all the scenarios are multiturn, and a judge grades the evaluatee on a scale from 1-5. Here, SteerDuplex is much closer to saturating the benchmark, especially on Safety.
Finally, VoiceBench is a single-turn benchmark that basically turns other existing benchmarks into voice versions. SteerDuplex mostly doesn’t improve here, with one outlier significantly impacting the overall score, so I won’t dwell on it.
By the way, they report later on how the model changed on the above metrics after RL. Basically, it didn’t, which really drives home how focused the RL was on timing specifically.
Now we can move on to RL results, which reflect behavior more than content.
On the left we see how often SteerDuplex barges in during a pause from the other party, like if you say “um” and think for a bit, intending to continue talking soon. The final version of SteerDuplex, after both stages of RL, seems to do much better at judging when to take the floor or not. Actually there’s some evidence from FDB v2 that the model errs the other way now, yielding the floor a little too readily.
On the right we have sort of the opposite case, where the user starts to overlap with the model. In the case of a genuine interruption, the model should yield, while in the other three cases the model should continue on. They graphed success rate, so higher is always better, i.e. a higher score on Interruption does not mean the model interrupted more.
Now we get to talk about the fun bit of RL: reward hacking.
Earlier I described the rewards they used in RL, each measuring different elements of performance. You might have wondered, how did they come to that exact formulation?
Well let me tell you, just looking at their list I could tell it was an empirical process, not a spec they designed once and then implemented. A lot of what gets rewarded and by how much comes from looking at model behavior at every stage of the training process and trying to compensate.
Sometimes you’re trying to compensate for new bad behavior or weaknesses you didn’t notice before, just adding to your reward picture. But other times you are reacting to behavior you yourself induced with your earlier rewards. In other words, you are responding to reward hacking.
The chart here is a good example. For some of the rewards in this paper, silence actually earned the model just as much as an appropriate response. For example, a reward for stopping a response when the user interjects is easy to reap if the model was already silent.
What the chart shows is how bad the “empty speech” problem gets when you isolate parts of the reward. Like if you look at the transcript only for example, half your outputs are empty. The full set of rewards mostly addresses the problem but doesn’t solve it.
Of course with more rewards there often comes a tradeoff. Here for example, the authors saw noise robustness decline while behaviors around interruption improved. You can adjust weights of the various rewards to reflect the balance you want to strike, but that introduces yet another thing to optimize and can get expensive.
My Takeaways
Audio still has a long way to go
Speech is more specifically human than writing
Full-duplex models can’t be that big due to latency. Clever engineering required to make them smart enough to talk to like a peer
Audio is a different world of research
Image input is common now - photos, screenshots - and required for a lot of important benchmarks
Audio is more about products than RSI
Concurrent interaction will require product changes
Text is easy because it is single-stream, you can’t really write or read at the same time as you do anything else
But it’s common to talk while doing something else, and we will expect that of agents
You’ll know audio has made it when there is a voice equivalent of AI slop















