GAME AUDIO
What a Game Audio Brief Is Actually Asking For
When a casting spec comes back from an audio director, most producers quietly search half the terms in it. Game casting runs on different logic to commercial or corporate work, and the reason is volume. Here is what a brief is really asking for, and why getting the words right is what gets you an accurate quote from gaming voice over services rather than a number that turns out to be wrong by a factor of three.
Structuring a Game Cast: Hero VO, NPC Banks, and Barks
Ask a studio how many characters are in their game and you will get a number. Ask what that number costs and the honest answer is that it depends entirely on the shape of the script, because a game is never cast as one block.
Hero VO is the part everyone pictures. The named, story-critical roles: playable leads, the antagonist, the companion who talks at you for forty hours. They carry the cinematic dialogue and take the most direction, and they are usually the smallest part of the bill.
The NPC bank is where the line count actually lives. Quest-givers, shopkeepers, guards, the nine-line character whose entire job is pointing you at a dungeon. One actor will often cover six of these in genuinely distinct voices, because nobody is paying separate day rates for a blacksmith and a market trader.
Then there are barks, which are short reactive lines triggered by something happening on screen rather than by a conversation. A guard clocking you. An enemy calling a flank. Someone muttering about the weather as you walk past. Barks live or die on repetition, and the moment a player hears the same one twice in a row it stops being world-building and starts being a joke. That is why bark sessions run far longer than the script length suggests: you are not recording forty lines, you are recording forty ideas eight ways each.
The one that quietly wrecks budgets is loop group, or walla. Background human sound rather than scripted dialogue: a market crowd, corridor chatter, a stadium reacting. It rarely appears in the first version of a brief and almost always appears in the final one.
Where the hours go
Count the bank, not the cast
Cost is driven by line count in the NPC bank and the number of bark variations, not by how many named characters are in the story. Ask for both figures before anyone quotes.
Booth Deliverables: Exertions, Wildlines, and ADR
This is where producers from other disciplines come unstuck, because several of these session types have no equivalent anywhere else in voice work.
Exertions, more commonly called efforts, are the non-verbal half of the job. Jumps, climbs, lifts, hits taken, falls, deaths, breathing under load. A combat system runs on them, and they are brutal on a voice in a way dialogue simply is not. Book a full day of deaths and you will not get usable hero dialogue out of that actor afterwards. Efforts get their own block, and they get scheduled last, never as a warm-up before a cinematic session.
Wildlines are recorded without picture or sync reference, outside the scripted flow. Nine times out of ten they are the lines design only realised it needed after a playtest.
Pickups are re-records of lines already in the can, because the script moved or the read did not sit right once it was in the build. Every game has them. A schedule without a pickups window is a schedule that has not met a playtest yet.
ADR means recording to picture, in sync with animation that already exists. In games it turns up mainly in cinematics and in localisation. Performance capture goes further, recording voice at the same time as body and face, so the performance arrives as one take rather than a voice grafted onto an animation weeks later.
Session tip
Never stack efforts on dialogue
Efforts wreck a voice faster than anything else in the booth. Give them their own block at the end of a run, not the morning before a cinematic session.
Engine Integration: Middleware, Visemes, and File Naming
This is the part that looks like admin and is actually the entire job. A mid-sized title runs to tens of thousands of individual audio files, and every one has to drop into the engine without a human opening it to check.
The naming convention is what makes that possible. Character, scene, line ID, take, specified by the audio team before recording starts. Get it wrong and it is not a small fix, it is somebody’s week spent renaming files by hand while the build waits.
Those files are then triggered through middleware, almost always Wwise or FMOD, with rules attached. It is the rules rather than the audio that stop a bark firing twice in a row. Visemes, the mouth shapes the animation is built from, cap how long a line can run, which becomes a serious constraint the moment you localise into a language that takes longer to say the same thing.
You do not need to operate any of this. You do need to hand over files that go straight in untouched, and that single habit is the clearest difference between a supplier an audio lead books again and one they do not.
The thing that gets you rebooked
Their convention, not yours
Ask for the naming spec before the first session, not after the first delivery. Renaming forty thousand files by hand is somebody’s week.
Game Localisation: LQA, Lip Sync, and Single-Pipeline Management
Localisation means recording the title again with native performers in each language. It is not a subtitle exercise, and treating it as one is where timelines fall apart. On a dialogue-heavy game it is usually the largest single line in the audio budget, and it needs the same casting discipline as the original. A hero voice in French has to land as a hero voice, not as a competent stand-in.
Two things decide whether it works. LQA, linguistic quality assurance, means checking recorded lines inside the actual build rather than on paper, because context breaks translations in ways a spreadsheet cannot show you. A line that reads perfectly in isolation can be wrong the moment you know who is saying it and who they are saying it to. And lip sync constraints mean a translation that runs long has nowhere to go when the face is already animated to a fixed length.
Sourcing compounds both. Eleven languages from eleven suppliers means eleven naming conventions arriving on eleven different days, each needing its own QA pass. Our multilingual voiceover and dubbing work runs through one pipeline for exactly that reason.
Quick gut check
One pipeline or eleven?
Every extra supplier is another naming convention, another delivery format and another QA pass. Consolidating is usually cheaper than the day rate suggests.
Future-Proofing Your Game Audio Assets
Game audio outlives the launch, which is what makes usage trickier than it first appears. Trailers, marketing cutdowns, DLC, a remaster three years out. Every one of those is a separate question, and every one is cheaper to settle at contract stage than to reopen later with an actor whose rate, availability or career has moved on.
Session structure follows the same logic. Game work is almost never a single booking. It is a block for hero VO, a bank for NPCs, a separate efforts day, pickups running through production, then a final pass at cert. Plan for that shape from the start and your casting stays consistent. Skip it and you find out eighteen months in that your lead has taken another job.
None of this vocabulary is gatekeeping. It exists because “some background chatter” is not a brief when the script runs to forty thousand lines, and “the guard shouting when he sees you” only becomes something anyone can quote once it is a bark with a defined trigger and a stated number of variations. Write the brief in these terms and the number that comes back will actually mean something.
Ready to brief your next game title? Listen to our video game voice actors or speak to a casting director on 0207 183 3750.

