Sierra's multimodal agents bring voice, text, and visuals into the same customer conversation. Voice is for explaining what you need, a visual for comparing options side by side, and text for referencing something later. Instead of picking just one, the agent automatically shifts between modes as the conversation needs.