← All posts

AI voice agents got a face and a 95% price cut

2026-09-26 · 4 min read

Two posts this week, from two companies on opposite sides of the Pacific, point the same way. Voice AI is getting cheaper to run, and it's getting a face.

On September 24, @GoogleCloudTech announced that Gemini 3.8 Live with Live Avatar is generally available in Gemini Enterprise, listing "video avatars, fluid dialogue, tool calling, and more." Around the same time, @Alibaba_Qwen introduced Qwen-Audio-3.1, describing it as "five models, one complete audio stack," and paired the launch with big price cuts across the lineup.

These are smaller launches than a new frontier model, but read together they tell you more about where customer-facing AI is headed.

What Google shipped

Live Avatar puts an animated person on screen on top of Gemini's real-time voice model. Google's announcement says the avatar lip-syncs, holds natural expressions, and takes turns in conversation across 97 languages. It can switch languages partway through a conversation without the lip-sync falling out of step.

For businesses, the important part is asynchronous tool calling. The avatar keeps talking while it does work in the background, like looking up a record or checking a policy. Google's examples are customer service, interactive walkthroughs, and hotel check-in. Coverage of the launch also describes an insurance claims demo: a customer holds damage up to the camera, and the agent checks coverage rules and assembles a packet for an adjuster while the call goes on.

There are limits. Companies can pick from preset avatars, but building one from a reference image, such as your own face or a spokesperson's, is restricted to allowlisted enterprise customers. The announcement doesn't list pricing. Every piece of audio and video carries Google's SynthID watermark, which is invisible to viewers and can be detected later.

What Alibaba shipped

The Qwen release covers the ears and the mouth. There are upgraded speech recognition (ASR), text-to-speech (TTS), and Realtime models, plus two new ones: ASR-Next, which separates multiple speakers and picks up tone and background noise, and TTS-Next, which generates voice, sound effects, and ambient audio in one pass.

The pricing gets the headlines. Reports on the launch put the cuts at up to 95 percent for speech recognition, about 70 percent for text-to-speech, and about 85 percent for the Realtime conversational model. The standard ASR model is reported to handle 30 languages and 16 Chinese dialects, and it removes filler words and repetitions from transcripts automatically.

Our read

The price cut matters more than the avatar, even though the avatar will get the screenshots.

The basic pieces of a voice agent are listening, thinking, and talking. Anything that takes an order, books an appointment, or answers "are you open Saturday" is built from those. The listening and talking parts are becoming a commodity. When a large provider cuts speech recognition prices by up to 95 percent in one announcement, everyone else feels pressure to follow. The cost of an AI phone line is moving toward the cost of the thinking in the middle, and that cost is also dropping.

We're more skeptical about the avatar. A talking face helps in some settings: a hotel kiosk, a guided product demo, a multilingual help desk. For most small businesses, customers want the problem solved, and a face on screen doesn't do that. They'd take a quick text reply or a phone call that books the job over an animated rep. The allowlist also means the most interesting option, an avatar that looks like your actual brand, isn't open to most companies yet.

Both launches handle the plumbing the same way, and that's the part worth copying. The voice or face is the front end. The value comes from the tools behind it: calendar, CRM, policy lookups, job notes. Google's background tool calling is really a claim about that back end.

What this means for a local business

If you've written off an AI receptionist or after-hours line as too expensive or too robotic, look again over the next few months. Prices are falling quickly, and response speed is getting close to human.

What decides whether it works for you is the same as last year. Your booking rules, service areas, pricing logic, and handoff to a human have to be written down clearly enough for software to follow. Most businesses we talk to in the Fox Valley haven't done that yet, and no model release will do it for them.

That's the job of our free process audit. We go through how calls, forms, and follow-ups actually move through your business, then show you which pieces are ready for a voice agent and which need fixing first. Picking the voice comes after that, and it's the easier decision.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere