Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two real-time audio models designed to hold natural conversations while reasoning, seeing visual context and using tools in the background.

For a small business, the important change is not a voice that sounds slightly more human. It is the ability to keep talking while checking an order, finding an appointment or completing another multi-step task instead of forcing the customer to wait in silence.

Google has also published usage-based API prices and a free testing route through AI Studio. That lowers the barrier to a prototype, but it does not make a reliable customer-service agent automatic. Phone costs, integrations, testing, privacy, human handover and mistakes can cost more than the model itself.

01

What Google launched

Confirmed fact: Google announced Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking on 15 September 2026 and updated the launch details on 17 September. The standard Live model is intended for fast, fluid conversations at scale; Extended Thinking is aimed at more difficult, multi-step tasks.

Both are audio-to-audio models available to developers through the Gemini API and Google AI Studio. Google says the models can process visual input in near real time, switch automatically between 97 supported languages and continue a conversation while tool or API calls run in the background.

Extended Thinking can reason and speak at the same time. It can acknowledge a request, narrate progress and continue the conversation while working through a more complex process. Google says all audio generated by its AI products contains an imperceptible SynthID watermark.

  • Near-real-time voice conversations
  • Automatic transitions between 97 supported languages
  • Visual context from a camera or shared screen
  • Background tool and API calls during the conversation
  • Extended Thinking for more complex multi-step work
02

Who can use it now?

Confirmed fact: developers can access both models through the Gemini API and try them in Google AI Studio. Gemini 3.8 Live is also rolling out in Search Live, while enterprise access is initially in private preview.

Google says Extended Thinking is rolling out in Gemini Live and to Google AI Pro and Ultra subscribers in Workspace Docs. It is also available in Gmail and Keep for Google AI subscribers. Availability may vary by account, product, country and rollout timing.

A business owner who wants a customer-facing agent will normally need more than a consumer subscription. The API provides the model, but a production service still needs a phone or web-audio layer, business data, tool connections, monitoring and a safe route to a person.

03

What does Gemini 3.8 Live cost?

Confirmed fact: Google's pricing page lists a free tier with limited access. On the paid tier, audio input costs $3 per million tokens, shown as approximately $0.005 per minute, and audio output costs $12 per million tokens, shown as approximately $0.018 per minute. Text, images, video, Google Search grounding and external services can add separate charges.

A simple example helps. If a customer speaks for three minutes and the AI replies for two minutes, the published model-audio charge would be about $0.051: $0.015 for input plus $0.036 for output. That is not the full cost of a five-minute customer call.

Telephony, orchestration software, storage, search requests, calendars, customer databases, developer time, tax and failed or repeated calls can all increase the total. Measure the completed outcome and the human correction time—not just the model's price per minute.

  • Audio input: about $0.005 per minute on the paid API tier
  • Audio output: about $0.018 per minute on the paid API tier
  • Three input minutes plus two output minutes: about $0.051 in model audio
  • Phone, tools, search, storage and integration costs are separate
  • Pricing and availability should be checked again before launch
04

Five practical small-business uses

The safest starting points are narrow, repetitive conversations where the correct answer already exists in approved business information. The agent should collect or retrieve facts, complete a limited action and hand over when the request falls outside its boundaries.

A trades business could qualify a new enquiry and offer available survey slots. An online shop could answer an order-status question after verifying the customer. A consultant could guide a new client through an onboarding checklist. A course creator could provide spoken navigation through approved lessons and frequently asked questions.

Visual context creates another useful option: a customer can show an item, screen or simple problem while speaking. That may help with product setup or troubleshooting, but it also raises privacy and accuracy risks, so the workflow needs explicit consent and strict limits.

  • Enquiry triage and appointment booking
  • Verified order-status updates
  • New-customer onboarding
  • Approved product or course guidance
  • Low-risk visual troubleshooting with human escalation
05

The opportunity: build one narrow voice workflow

Our analysis: freelancers and small agencies can package voice-agent discovery and prototyping as a service. The useful offer is not 'replace your receptionist'. It is 'test whether one repeated conversation can be handled accurately, safely and at a sensible cost'.

A starter package could map one call type, prepare an approved knowledge file, define tool permissions, write the handover rules, test 25 realistic conversations and provide a cost-and-error report. The customer keeps control of the phone number, accounts and final launch decision.

Choose a niche you understand—such as trades, property enquiries, training providers or ecommerce support—and solve one bounded problem. The advantage comes from the workflow, testing and knowledge of the business, not from reselling access to the same model everyone can reach.

06

The risks and limitations

A natural voice can make an incorrect answer sound more trustworthy. The agent may misunderstand a name, invent a policy, expose information to the wrong caller or complete an action the customer did not intend. Background tool use increases the damage a mistake can cause.

UK businesses must handle personal data lawfully and explain recording, transcription or automated processing where required. Sensitive identity, payment, health, legal or employment decisions need stronger controls and often a qualified person. Do not copy customer calls into a free developer account without checking the applicable data terms.

A production agent needs identity checks, least-privilege permissions, spending and action limits, logs, red-team tests and an obvious way to reach a human. It should never conceal that it is AI, pressure callers to continue or pretend to have completed a task when a tool failed.

07

One practical action: run a 25-call shadow test

Pick one low-risk call type that currently follows a clear script. Write the exact facts the agent may use, the two or three actions it may take and every condition that must trigger a human handover. Remove payment taking and other irreversible actions from the first test.

Build a private prototype in Google AI Studio or with a developer, then run 25 scripted conversations without connecting it to real customers. Include accents, background noise, interruptions, vague requests, angry callers, unavailable appointments and attempts to obtain private information.

Score each test for correct understanding, correct facts, correct action, handover quality, response time and estimated total cost. Only consider a small live pilot when every serious failure is fixed, the business can monitor calls and a person can take over immediately.

  • Choose one low-risk call type
  • Limit the approved information and available actions
  • Write mandatory human-handover conditions
  • Test 25 normal and difficult conversations privately
  • Launch only after serious failures are fixed and monitoring works
FAQ

COMMON BEGINNER QUESTIONS

What is Gemini 3.8 Live?

It is Google's near-real-time audio model for natural voice conversations, visual context and background tool use. An Extended Thinking version is designed for more complex multi-step tasks.

How much does Gemini 3.8 Live cost?

Google currently lists paid API audio input at about $0.005 per minute and audio output at about $0.018 per minute. Telephony, tools, search, storage and development are separate costs.

Can I try Gemini 3.8 Live free?

Google lists a limited free API tier and says developers can try the models in Google AI Studio. Limits and data-use terms differ from paid production access.

Can Gemini 3.8 Live replace a receptionist?

It can support narrow tasks, but a safe deployment still needs approved information, identity checks, limited permissions, monitoring and immediate human handover.

What should a small business test first?

Start with one low-risk conversation such as enquiry triage or appointment availability, then run at least 25 private test calls before involving real customers.