Sierra's multimodal agents are AI agents built for customer conversations that bring voice, text, and visuals into the same interaction. Instead of forcing a customer to choose a single medium, the agent automatically shifts between modes as the conversation requires. Voice is used when a customer wants to explain what they need, a visual when it helps to compare options side by side, and text when someone wants to reference something later. Sierra frames the result as an interface that morphs with the conversation, so customers get the best of each medium without having to pick just one. The agents are intended for companies that handle customer interactions and want those interactions to feel continuous rather than fragmented.
The problem these agents address is familiar to anyone who has tried to complete a purchase or a change over the phone. Sierra uses the example of upgrading a mobile plan: a representative talks through models, colors, storage sizes, and monthly rates, and the customer is left comparing all of it in their head and picking a phone they cannot picture. The call is genuinely good for parts of the task, because it is easier to say what you actually need and to ask questions than it is over text. But the customer cannot see the thing they are about to buy, and that gap makes the decision harder and slower than it needs to be. Sierra's multimodal agents are described as closing that gap by bringing voice, text, and visuals into the same conversation.
Connecting multiple channels is not, in Sierra's framing, the difficult part. The real trick is knowing which modality to use when: voice to explain what you need, a visual to compare options side by side, or text when you want to reference something later. Agents built on Sierra anticipate what is needed for each conversation and automatically shift between modes. Crucially, that switching happens without making the customer start over or repeat themselves, which is what usually happens when an interaction moves between a phone call, a chat window, and a self-service screen. The interface is meant to follow the conversation rather than the other way around, so the customer never has to re-explain context that the agent already has.
Concrete examples show how that plays out in practice. If a flight is disrupted and a customer calls the airline to get a new flight, instead of a representative reading alternate options off one by one, the options appear laid out with departure times, layovers, and pricing right in the conversation. The customer picks one, and the agent keeps going from there. Choosing a seat works the same way: the customer sees the seat map and taps the seat they want. And for times when it is easier to talk than to type, the customer can switch to voice and explain exactly what they need, with the agent capturing those details instead of asking the person to type a paragraph into a text box.
Multimodal agents also follow Sierra's approach of one agent for every surface. You can build your agent once and deploy it across all channels, and the same is true of multimodal components: once you build a visual component, your agent can use it everywhere it lives. That approach extends to Sierra's MCP UI integration, which lets you bring interactive components such as product cards, comparison tables, calendars, and forms directly into the conversation. Your team designs and hosts those components, so you decide how they look, what they show, and when they change. When you make an update, it is automatically reflected everywhere without needing to redeploy or maintain different versions for each platform.
When something needs more room, a component can expand to full screen to show calendars, long comparison tables, multi-step forms, and more. This matters because the interface can scale with the complexity of a task without the customer ever leaving the conversation. The underlying idea is that the conversation is the interface: the customer says what they need and the agent figures out the rest, using whichever mode of communication fits the moment. Multimodality is the mechanism that lets that principle hold true across tasks that involve talking, reading, comparing, and tapping to confirm.
For customers, the outcome is that they do not have to choose. On a single call, a customer can talk through what they need, glance at a screen to compare their options, and tap to confirm, without ever pausing the conversation to switch tools. That continuity removes the repetition and dead ends that usually come with moving between a phone call and a screen. For the businesses deploying these agents, the stated benefit is that multimodal experiences are as easy to build and deploy as they are for customers to use, thanks to the build-once, deploy-everywhere model and components that are owned and updated centrally.
Use cases described in the content include telecommunications journeys such as upgrading a mobile plan, where the customer talks through models, colors, storage sizes, and monthly rates while seeing the options rather than holding them in their head. Travel is another: rebooking a disrupted flight while alternate options with departure times, layovers, and pricing appear in the conversation, and selecting a seat by tapping a seat map. Forms, calendars, product cards, and comparison tables can all be embedded where a conversation needs them, including multi-step forms that can expand to full screen.
The agents are aimed at organizations that run customer interactions and want to design the experience themselves. Sierra emphasizes that your team designs and hosts the components used in the conversation, which gives you control over how they look, what they show, and when they change. Deployment is described as building the agent once and running it across all channels, with updates flowing everywhere without redeployment or platform-specific versions. The MCP UI integration is the specific integration named in the content for bringing interactive components such as product cards, comparison tables, calendars, and forms directly into the conversation.
The takeaway is that Sierra's multimodal agents treat the conversation itself as the interface and let the medium change as the moment demands. Voice handles explanation, visuals handle comparison, and text handles referencing, all within one continuous interaction that customers never have to restart. Built once and deployable everywhere, with components your team owns and updates centrally, the agents aim to make multimodal customer experiences as straightforward to build as they are to use.