Waterr AI Logo
ResearchJuly 1, 2025

Video AI Agents: Why the Video Call Is the Next Frontier

A video AI agent isn't a chatbot with a face — it's an AI that sees, listens, and responds inside a real video call. Text gives AI words. Audio gives it tone. Video gives it everything — expression, attention, body language, environment. The companies that ship video AI agents will define the next platform.

Harshit SharmaFounder & CEO, Waterr AI

Here is a claim I'll defend from first principles: the video call is the highest-bandwidth interface humans have ever built for transferring context between two minds. And any AI system that ignores this channel is leaving most of the signal on the table.

Video AI agents are AI systems that participate in live video calls — not as a bot in the corner writing notes, but as a full participant that sees your face, hears your voice, tracks your screen, and responds in real time. They are the next form factor for AI, and they exist for the same reason spreadsheets and search boxes did before them: they move AI onto the channel where the actual signal lives.

Text gives AI words. Audio gives it tone. Video gives it everything — expression, attention, body language, environment. Below is the first-principles case for why a video AI agent is the highest-bandwidth interface an AI can have with a person, and what a stack has to look like to actually build one.

Let me make the case.

The Bandwidth Argument

When you type a message to ChatGPT, you transmit roughly 40–60 bits per second of information. That's the rate of human typing — about 40 words per minute, 5 characters per word, 1–2 bits of entropy per character.

When you speak to a voice assistant, you transmit around 2,000–4,000 bits per second. That includes not just words, but prosody — pitch, rhythm, emphasis, hesitation, sarcasm, excitement. A 50x improvement over text.

When you sit in a video call, the bandwidth explodes. The camera alone transmits millions of bits per second — your facial expression, your eye contact, your posture, your gestures, the objects on your desk, the room behind you, whether you're leaning in or checking your phone. Add the audio channel, and you have a multi-modal data stream that is orders of magnitude richer than anything text or voice can provide.

This isn't an aesthetic preference. It's information theory. Claude Shannon proved in 1948 that the value of a communication channel is determined by its capacity. A video call has more capacity than any other human-computer interface that exists today. Every AI system that limits itself to text is voluntarily operating on a fraction of one percent of the available signal.

The question isn't whether video will become an AI interface. It's why it took this long.


What a Camera Sees That a Chatbox Can't

Consider what a single frame of video contains that text never will:

  • Facial expression — confusion, understanding, skepticism, excitement, frustration. Humans communicate more through micro-expressions than through words. Paul Ekman catalogued over 10,000 distinct facial configurations. A text prompt captures none of them.
  • Attention and engagement — is the person looking at the camera? At their phone? Are they leaning forward or slouching back? Are they nodding along or staring blankly? This tells you whether your message is landing — information that is invisible in a chat interface.
  • Body language — arms crossed (defensive), hands gesturing (animated), head tilted (curious). Albert Mehrabian's research suggested that up to 55% of emotional communication is visual. Even if you discount the exact number, the direction is unambiguous: most human communication is non-verbal.
  • Environment — a messy desk, a hospital room, a conference room with five people, a living room with a child in the background. The setting provides context that the user would never think to type out. It tells you whether this is a casual conversation or a high-stakes meeting before a word is spoken.
  • Screen content — when someone shares their screen, the AI sees exactly what they see. Not a description of the problem. Not a screenshot they remembered to take. The actual, live, pixel-perfect state of their work.

Now imagine an AI that processes all of this — not sequentially, not after the call, but in real time, during the conversation. That's not a chatbot with a camera. That's a fundamentally different category of AI system.


Reading the Room Is a Computational Problem

Humans are extraordinary at reading rooms. You walk into a meeting and within three seconds you know: the energy is tense, the CEO is frustrated, the presenter is nervous, two people in the back aren't paying attention. You don't consciously analyze this. Your brain processes it as a gestalt — a unified perception synthesized from hundreds of visual and auditory signals.

For decades, this was considered impossible for machines. Not because of compute — because of representation. How do you encode "the room feels tense" as a loss function?

Large multimodal models changed this. A vision-language model doesn't classify emotions into six buckets. It describes what it sees in natural language — "the person looks uncertain, their brow is furrowed, they're looking away from the camera" — and passes that description to a reasoning model that can act on it. This is holistic scene understanding, not reductive emotion classification.

In our architecture, this happens every few seconds. A vision model analyzes the camera frame and injects a natural-language description into the conversational context. The main LLM doesn't see pixels. It sees: "The participant appears to be thinking deeply, looking down, hand on chin. There's a whiteboard behind them with a diagram that looks like a system architecture."

The LLM can now make decisions that a text-only system never could:

  • Pause because the person is still processing
  • Ask a clarifying question because the expression suggests confusion
  • Reference the diagram on the whiteboard behind them
  • Adjust its pace because the person seems disengaged

This is reading the room. And it's now a computational problem with a working solution.


Why Meetings Are the Richest Data Source for AI

There's a reason humans default to meetings for anything important. Not email. Not Slack. Not documents. Meetings.

A 30-minute meeting transfers more context than 50 emails. That's not hyperbole — it's a consequence of bandwidth. In 30 minutes of face-to-face conversation, you transmit:

  • The explicit content — what you actually said
  • The implicit content — how you said it, what you emphasized, what you skipped
  • The reactive content — how the other person responded to each point, what made them lean in, what made them look away
  • The relational content — trust signals, rapport, shared context, inside references

This is why you can't onboard someone over email. Why you can't close a deal over Slack. Why you can't resolve a conflict over a document. The bandwidth isn't high enough.

Now consider AI as a meeting participant. Unlike a human, AI can absorb every channel simultaneously — audio, video, screen share, chat, and transcript — without cognitive fatigue. It doesn't zone out. It doesn't miss the micro-expression. It doesn't forget what was said 20 minutes ago. It processes the full bandwidth, all the time.

This makes meetings the single richest data source for building AI that truly understands human context. Richer than browsing history. Richer than email archives. Richer than any text corpus ever assembled. Because meetings are where humans are most fully themselves — communicating with every channel they have.


The Dermatologist, the Interviewer, the Advisor

Once you accept that video is a first-class AI input, use cases emerge that are simply impossible with text:

Visual assessment. A dermatologist needs to see your skin. A physical therapist needs to see your posture. A fitness coach needs to see your form. Video makes the AI's eyes as useful as its language. Not as a replacement for a doctor — but as a first-pass screening that's available 24/7, in any language, at zero marginal cost.

Behavioral assessment. An interviewer doesn't just evaluate answers — they evaluate how you answer. Confidence, hesitation, clarity under pressure, how you handle a curveball. An AI interview coach with video can give you real feedback on your body language, eye contact, and pace — not just your words.

Adaptive coaching. A sales trainer that watches your pitch practice and notices you speed up when you get to pricing — a sign of discomfort. A presentation coach that sees you reading from notes instead of making eye contact. A language tutor that watches your mouth shape to correct pronunciation. All impossible without video.

Real-time advisory. You're wrestling with a complex decision. You don't want to write a 2,000-word prompt. You want to talk it through — the way you would with a trusted advisor. You want the AI to see your screen when you pull up the spreadsheet. You want it to notice when you're unconvinced by its suggestion. Video makes this feel like a conversation between colleagues, not a form submission.

This is what Plus One Studio enables. You create an AI persona — with a specific expertise, demeanor, voice, and set of goals — and deploy it as a live video meeting agent. An interviewer persona that's analytically rigorous. A sales coach that's blunt but supportive. An onboarding guide that's patient and friendly. Each one configurable, each one seeing and hearing the full context of the conversation.


Beyond Book-a-Demo Buttons

Here's a thought experiment for every SaaS founder:

Your website has a "Book a Demo" button. A potential customer clicks it. They fill out a form. They wait 24 hours. A sales rep emails them. They schedule a call for next week. The rep joins the call, asks the same qualifying questions the form already asked, and finally — seven days after the initial interest — the customer gets to explain what they actually need.

Now imagine this: the customer clicks the button and is immediately in a video call with an AI that has read your documentation, understands your product, and can answer questions in real time. The AI sees the customer's screen if they share it. It adapts its tone — technical for engineers, strategic for executives, patient for first-time buyers. It qualifies the lead. It can schedule a follow-up with a human rep if needed. Zero latency between intent and conversation.

This isn't a chatbot with a video feed. This is a paradigm shift in go-to-market. The conversion funnel collapses because the highest-bandwidth communication channel — video — is now available at the top of the funnel, not reserved for the bottom.

And the AI persona doesn't take lunch breaks, doesn't have a quota, doesn't need to check with a manager. It's available in every timezone, in every language, for every visitor, simultaneously.


The Architecture of Real-Time Understanding

Building video-native AI isn't just plugging a camera into a chatbot. It requires a fundamentally different architecture.

The pipeline is bidirectional and real-time. Audio flows in through WebRTC, gets transcribed by a speech-to-text engine in milliseconds, feeds into a large language model that reasons about the conversation, generates a response, passes it to a text-to-speech engine, and plays it back — all while the video stream is being analyzed in parallel by a vision model. This entire loop runs in real time, with voice activity detection deciding when the AI should speak and when it should listen.

Multimodal fusion happens at the context level. The vision model doesn't control the conversation. It provides context. Every few seconds, it describes what it sees — "the participant is smiling and nodding" or "they're looking at a chart on their screen" — and injects that into the LLM's context window. The LLM fuses audio understanding (from the transcript) with visual understanding (from the vision descriptions) and generates a response that accounts for both.

Memory persists across sessions. The AI remembers what happened in the last meeting with this person. What was discussed, what was decided, what follow-ups were promised. This transforms the AI from a single-session tool into an ongoing relationship.

Personas are configuration, not code. The same architecture powers a gentle onboarding guide and a rigorous technical interviewer. The difference is the system prompt, the demeanor parameter, the voice selection, and the scenario goals. Creators define these in a UI. No engineering required.

Meeting detection is ambient. On the desktop, the system monitors your operating system's audio stack. When your microphone activates for more than five seconds — indicating a meeting has started — recording and transcription begin automatically. You don't press a button. The harness is always ready.


AI as Your Brainstorming Partner

There's a deeper argument for video AI that goes beyond features and use cases. It's about how humans think.

We don't think by reading. We think by talking. Dialogue is the original reasoning engine — Socrates knew this 2,400 years ago. When you explain a problem out loud, you understand it better. When someone asks you a question you didn't expect, you discover what you actually believe. When you see confusion on someone's face, you realize your explanation wasn't clear.

Reading is serial and passive. Conversation is parallel and interactive. The feedback loop is tighter. The bandwidth is higher. The processing is faster — not because the information density is higher, but because the human brain is optimized for face-to-face communication. We evolved for it over millions of years. We've been reading for barely five thousand.

AI is extraordinarily good at reading. It can process a 100,000-token document in seconds. But you can't. You need time. You need to process. You need to react, question, push back, reconsider.

A video AI meets you where your brain actually works best — in conversation. It reads the document for you, synthesizes it, and discusses it with you while watching your face to see if you're following. It adjusts. It slows down. It notices when you're lost. It speeds up when you're bored.

This isn't a convenience feature. It's a cognitive amplifier. The AI handles the bandwidth it's good at (reading, synthesizing, recalling) and communicates through the channel you're good at (conversation, visual processing, intuitive judgment).


The Next Interface

Text was the first AI interface. It gave us ChatGPT and a revolution in accessibility. Anyone could prompt a model.

Voice was the second. It gave us Siri, Alexa, and eventually real-time voice agents. The bandwidth went up 50x.

Video is the third. And the bandwidth increase isn't 50x. It's 1,000x or more. Facial expressions, body language, attention signals, environmental context, screen content, real-time visual feedback — all flowing into an AI system that can reason about every channel simultaneously.

Every previous AI interface was a compression of human communication. You compressed your thoughts into text. You compressed your intent into voice. Video is the first interface that approaches the full bandwidth of human expression.

The companies that build video-native AI — not video as an afterthought, but video as the primary input modality — will define how humans and AI work together for the next decade. Not because video is trendy. Because information theory demands it.

The richest signal wins. And video is the richest signal we have.

Video AI AgentVideo AIMultimodalPlus One StudioAI ResearchWebRTCComputer VisionFuture of AI