← All posts

Gemini 3.8 Live's budget model scored 30% on customer-service tasks

2026-09-16 · 4 min read

Google shipped two new voice models on Tuesday. @GoogleDeepMind introduced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking as its best conversational AI, models that "talk, think, and handle tasks in the background without breaking your flow." Google's @OfficialLoganK said they come with "frontier price + performance," and noted that 3.8 Live handles 97 languages, switches between them mid-conversation, and can run tool calls asynchronously.

About four and a half hours later, the benchmarking firm @ArtificialAnlys posted its results. Extended Thinking, run at high reasoning effort, debuted at No. 1 on its Speech to Speech Index with 82.6, just ahead of OpenAI's GPT-Live-1 at 81.5 and xAI's Grok Voice Think Fast 2.0 at 81.3. The standard 3.8 Live, at $0.84 per hour of input audio, is "the cheapest model in the Index."

That's what most people will remember. I'd look further down the same post, though.

What Google actually shipped

The feature a business should care about is the asynchronous tool calls. The model can look something up or make a booking in the background and keep talking while it waits, much like a receptionist who says "one moment" while checking the calendar. Google's announcement describes the Extended Thinking model saying things like "Let me check that" and narrating its progress on multi-step jobs.

The two models have different jobs. The standard 3.8 Live is Google's cost-efficient option for scale, and it now powers Search Live in the Google app. Extended Thinking is meant for harder work: it's rolling out to Gemini Live and runs the new voice features in Gmail, Docs, and Keep for AI Pro and Ultra subscribers. Developers can use both through the Gemini API and Google AI Studio, and Gemini Enterprise has them in private preview.

The number under the headline

Artificial Analysis also runs its own version of Tau Voice, a customer-service benchmark from the AI company Sierra. A simulated customer calls in with a problem, such as a flight change or a disputed charge. The model gets the company's policy documents and a set of tools, and it only passes if the records at the end match the correct outcome.

Extended Thinking scored 68.6% on that test, the best result. The standard 3.8 Live scored 30.1%.

Put another way, the budget model costs about a quarter of what Extended Thinking does (Artificial Analysis prices that one at $3.50 an hour), and it passed fewer than half as many tasks.

The standard model isn't weak across the board. In Artificial Analysis's Speech Agent Arena, real people try the same scenario (booking an appointment, for example) with two unnamed systems and pick the one they like better. The standard 3.8 Live ranked second on preference, ahead of Extended Thinking, and finished 93.2% of those arena tasks. It also starts talking sooner. Where it falls off is the longer call, where the agent has to follow a policy through several steps and leave the records right at the end.

Even the top model misses a lot. Google's own announcement cites 35.1% for Extended Thinking on Sierra's banking version of the test, and Sierra says the best text models reach about 85% on the original tasks. Voice agents are still well behind that.

My read

The progress is real. In the same comparison, the older Gemini 3.1 Flash Live cost $1.75 an hour and took 2.99 seconds to start speaking, while the new standard model costs about half that and starts in 1.18 seconds. Google also puts an inaudible SynthID watermark on all of the generated audio, which I'd like to see become the default everywhere.

My problem is with the pitch. It pairs the best benchmark with the lowest price, and those numbers belong to different models. If a vendor offers you an AI receptionist for "under a dollar an hour," ask which model is behind it. That answer tells you whether the bot will just sound good on the phone or can actually move the appointment.

What this means if AI answers your phone

Plenty of small-business calls are simple, like someone asking when you close or whether you can fit them in Thursday. A cheap, fast model may handle those fine. Rescheduling or applying a cancellation policy takes more steps, and those are the tasks where the scores drop. A sensible setup sends the easy calls to the cheap model and the multi-step ones to the stronger model, with anything unclear going to a person.

Before any bot answers your phone, test it on 20 of your real calls and count how many it gets right.

New Face Design's free process audit starts from the same place: we go through the calls and requests your business gets and sort out which ones software can finish on its own and which still need a person. Start here.

08 / Start here

Find your worst bottleneck. Free.

A 20 minute call. We map where your week goes and pick out the first process worth automating. You keep the map either way, and there is no deck to sit through at the end.

Email

pgorski@newfacedesign.com

Phone

+1 (773) 627-2176

Based in

Chicago area

Working with clients everywhere