Why a minutes forecast picks the wrong mode
The usual starting point is a spreadsheet: expected calls a month, average length, a mode chosen to fit. That answers a capacity question. It says nothing about whether a particular conversation can work in that mode, because an interpreter can interpret only what reaches them.
Take three encounters that could sit in the same monthly total (examples, not measured cases). A caller moving an appointment. A pharmacist explaining a dosing schedule with the bottle in hand. A field technician describing a fault on a component they could simply point at. The first works on a clean phone line. The second and third lose part of their meaning the moment the picture is gone, and no amount of volume planning puts it back.
Start with what has to be seen and heard
For each encounter type, write down what the participants refer to while they talk: a form, an image, a device, a room layout, a gesture. If the meaning depends on it, the interpreter needs to see it, or someone has to describe it, and describing is slower and less exact than showing. Reading a letter aloud line by line over the phone is a poor substitute for putting the same approved page in front of everyone.
Then look at the speakers. How many will talk, and can an interpreter working on audio alone tell them apart? Will a relative or a colleague cut in? A two-person, spoken-language exchange with a simple structure is where phone interpretation is strongest: quick to join, little equipment, nothing to frame. Add a third voice on a speakerphone at the far end of a desk, or pass one handset back and forth, and turn-taking starts to break down.
One case is not a trade-off at all. When the person receiving the service uses a signed language, an audio line is not an option, because the language itself is visual. The US Department of Justice's ADA guidance on effective communication sets out what video remote interpreting has to deliver there: real-time, full-motion video and audio, a picture large enough to show the interpreter's face, arms, hands and fingers, clearly audible voices, and staff trained to set it up quickly. Whether that guidance applies to a given organisation, and which aid is effective for a particular person, is for that organisation's own compliance or legal owner to decide, ideally after asking the person what works for them.
Check the setting and both ends of the connection
The setting decides what a mode can deliver. A bedside, a front desk, a home visit, a site office and a vehicle each strain a different part of a remote connection. Ask who else can hear the speakerphone or see the screen, whether the participant can move somewhere quieter, and whether the room allows a device at all.
Then check the equipment at each end, because there are usually more ends than there appear to be. When staff and participant share a room, one device serves two people: its microphone has to pick up both voices and its camera has to frame both faces, not the ceiling or a bright window behind them. When the participant joins from their own phone, their connection is the weak link and nobody on the buyer's side controls it. On the interpreter's side, confirm that the meeting platform the organisation uses is one the interpreter can join, and that a security setting will not block it on the day.
In a spoken-language session, a frozen picture delivers less than a clean audio line. Test the video setup in the actual room, on the actual device and platform, before the first live session, not in a demonstration on office Wi-Fi. The conditions in the ADA guidance make a sensible practical test for any video setup, signed or spoken: motion that keeps up, a picture sharp and large enough to read faces and hands, clear sound, and staff who know how to start it.
Healthcare and court settings may carry their own rules on when remote interpretation is allowed and what it has to include. Confirm them with the compliance owner, or with the court, before a mode is fixed. This guide is a planning aid, not a compliance ruling.
Write the fallback before the session
Every mode fails sometimes. What matters is the next minute. If the video drops, does the session continue on audio, move to a phone line, pause for a reconnect, or get rebooked with an interpreter in person? Decide it per encounter type, write it where staff will see it, and name the role that makes the call while the participant is waiting.
The fallback has to survive the failure it covers. Dropping from video to audio on the same weak connection is not a fallback. For encounters where audio cannot carry the meaning, such as a signed-language session or one built around a document everyone must see, an audio fallback is not one either; the realistic options are reconnecting, rescheduling or an in-person interpreter. If a mode has no workable fallback for an encounter type, reconsider the mode for that encounter.
Only then let volume shape the operation
Once each encounter type has a mode and a fallback, volume becomes useful. Count sessions by encounter type, not total minutes. A front desk handling many short spoken calls and a few long sessions around shared documents is running two setups, and an average of the two describes neither.
Volume then decides how each setup runs: booked or on demand, how the interpreter receives context before a long session, where devices are kept and who checks them. Track mode switches and abandoned sessions per encounter type as well. A minute total can rise while a mode is failing; a rising switch rate shows it.
Where the mode decision leads next
These pages pick up once the mode question is answered, or when one mode needs a closer look.
- Setting up an interpretation program (OPI and VRI): Use once the mode is chosen and the wider setup, from interpreter assessment to scheduling, needs designing.
- Over-the-phone interpretation: Use when encounters are spoken, short and carry nothing visual.
- Video remote interpreting: Use when a document, a device or a signed language is part of the encounter.
- Interpretation provider buyer guide: Use when the mode is settled and choosing the provider is the next decision.
Decide the mode for each encounter type
Work through this once per encounter type, not once for the whole service.
- The language, and whether the person receiving the service uses a spoken or signed language, confirmed with them rather than inferred from nationality
- What must be seen for the meaning to land: documents, images, devices, the room, gestures, faces
- How many people will speak, and whether an interpreter on audio alone could tell them apart
- The physical setting, and who else can hear or see the session
- The device, platform, camera position and connection at each end, tested where the session will happen
- The fallback if the chosen mode fails, and the staff role who calls it
- Sessions counted by encounter type before any monthly total is compared
Stop and rethink the mode
Any one of these means the mode was chosen before the encounter was understood.
- The mode was picked from a minutes forecast with no encounter types listed
- A signed-language participant is routed to an audio line, or offered one as the fallback
- Video is planned for a room nobody has tested for camera angle, light or connection
- Documents are read aloud over the phone when a shared page was possible
- The written fallback relies on the same connection that just failed, or there is none
- One mode is rolled out to every department because a single pilot went well
- A healthcare or court rule is assumed rather than confirmed with the compliance owner
What to send MoniSa
Send the encounters, not only the minutes, so the scoping conversation starts where the mode decision does. The buyer's own compliance owner still decides which sector rules apply.
- The language and dialect, and whether the person receiving the service uses a spoken or signed language
- Each encounter type, its typical length, and what participants need to see during it
- Where each end connects from, with the device and meeting platform in use
- Who joins, and any privacy or recording restrictions in the setting
- The fallback already in use, and expected sessions per encounter type each month