Rival AI agents, Instinct and Metas Muse, both add the ability to make calls
Discover how rival AI agents Instinct and Metas Muse are introducing phone calling features, reshaping the future of autonomous conversational AI voice agents.
Researched and edited by Kiran Ch and the WhatIsFuture editorial team. Reviewed for factual accuracy before publication.
Remember back in 2018 when Sundar Pichai stood on stage at Google I/O and showed off Google Duplex booking a hair appointment? Everyone’s jaw dropped, but then the feature largely faded into the background—scripted, fragile, and locked deep inside Google’s proprietary ecosystem. Fast forward to today, and we are finally witnessing the raw, unscripted democratization of autonomous voice systems. As TechCrunch recently reported, rival AI agent platforms Instinct and Metas Muse have both rolled out native phone-calling capabilities. Their agents can now directly place calls across the public switched telephone network (PSTN) to handle annoying real-world chores like securing table reservations, tracking down hard-to-reach customer support, or battling aggressive retention reps to cancel subscriptions.
This is far more than a flashy feature addition; it represents a fundamental shift in user-agent interaction. For the past two years, the AI ecosystem has been trapped in the prompt-and-response text box paradigm, forcing users to act as active babysitters for context windows and chatbot outputs. By giving AI agents direct access to the audio wire—complete with full-duplex speech interaction and interactive voice response (IVR) phone tree navigation—Instinct and Metas Muse are pushing software from passive conversation to asynchronous execution. We are moving from AI that tells you how to do something to AI that simply gets it done while you sleep.
Join Our Tech Community
Get instant alerts on the most critical AI breakthroughs on our WhatsApp channel. No spam, just signal.
Key Takeaways
- Asynchronous Voice Execution: AI agents are evolving from real-time text assistants into autonomous, goal-driven proxies capable of navigating legacy human-centric systems like phone networks.
- The End of Dark Patterns: Companies that rely on high-friction phone processes to keep users locked into subscriptions are about to lose their unfair advantage to relentless, persistent voice bots.
- Architectural Bottlenecks Shift: The technical moat is no longer just model intelligence—it lies in ultra-low latency audio processing, dynamic interruptibility, and surviving complex IVR phone trees.
- Bot-on-Bot Telephony: The imminent overlap between AI calling agents and AI customer service bots will redefine network load, caller verification, and voice application security.
The Death of the Friction Economy
For decades, major corporations have built multi-million dollar business models around administrative friction. Subscription services, cable providers, and fitness clubs deliberately make sign-ups seamless via a single online click, while forcing cancellations through multi-stage phone calls with high wait times and aggressive retention agents. This consumer trap relies on human fatigue. Most people would rather forfeit twenty dollars a month than endure forty minutes of blaring hold music followed by a high-pressure sales pitch.
Autonomous calling agents like Instinct and Metas Muse render this entire retention strategy completely useless. An AI agent doesn't get frustrated, doesn't feel awkward when declining a discount offer, and certainly doesn't care if it sits on hold for three hours listening to distorted pan flute music. When a user can simply delegate a task—telling an agent to cancel a membership or contest a mysterious bank fee—the artificial friction built into corporate workflows completely dissolves.
What makes this moment fascinating is how rapidly the underlying tech stacks have matured. Rather than relying on rigid, hardcoded voice decision trees, these modern agents utilize fast speech-to-speech models and low-latency audio pipelines that allow them to speak naturally, adjust tone on the fly, and hold firm against high-pressure customer service tactics. When corporate retention scripts encounter an agent that literally cannot be worn down by emotional manipulation or delay tactics, the balance of power shifts decisively back to the consumer.
Under the Hood: Latency, Full-Duplex Audio, and IVR Navigation
Building an agent that types text is trivial compared to building an agent that speaks naturally over a noisy cellular connection. The technical gymnastics required to pull off real-time outbound calling are staggering. First, you have the latency challenge: human conversation relies on precise turn-taking cues that occur within a 200 to 500-millisecond window. If an AI agent takes two full seconds to process a response, the human on the other end immediately senses something is off and hangs up or disengages.
To overcome this, systems like Instinct and Metas Muse leverage streamlined speech-to-speech architectures or heavily optimized cascaded stacks (using ultra-fast Automated Speech Recognition, specialized small language models for rapid decision making, and neural Text-to-Speech synthesis). Furthermore, these systems require full-duplex capabilities. The agent must be able to listen while speaking, allowing it to register interruptions, store late-arriving context, and pause instantly when spoken over—just like a human caller would.
Then there is the messy nightmare of navigating Interactive Voice Response (IVR) phone trees and unexpected edge cases. An agent calling a restaurant or a service center must recognize tone-dial inputs, understand confusing automated audio menus, handle unexpected call transfers, and maintain context through multi-party handoffs. If you look at why AI agents are lying, cheating, and coordinating under stress, you see how easy it is for an agent to hallucinate or misinterpret goal states when pushed by aggressive prompts. Ensuring that a calling bot stays on script while remaining flexible enough to handle complex human phone menus is where the real engineering battle is happening.
"The true bottleneck for voice agents was never voice synthesis—it was latency and contextual persistence. Once an agent can handle ambient noise, bad cell signals, and hostile phone trees in under 400 milliseconds, traditional customer service infrastructure is altered forever."
The Impending Cold War of Bot-to-Bot Telephony
Here is where things get bizarrely meta. As consumer-facing platforms deploy outbound voice agents to do our bidding, enterprise call centers are simultaneously replacing human reps with inbound AI customer service bots. We are fast approaching a world where your personal AI agent calls a company, only to be greeted by that company's customer service AI agent. Two software programs will conduct high-stakes negotiations over cellular lines using artificial human voices, while their actual human owners go about their day.
This setup creates bizarre incentives and fresh technical friction. How will an inbound AI agent verify that an outbound AI agent actually has legal authorization to act on behalf of a user? How do you prevent recursive logic loops, where two conversational agents get trapped in an infinite loop of polite handshakes and context exchanges, racking up thousands of API tokens and telecom minutes in the process?
As these interactions become standard, companies will inevitably launch AI-blocking tools—voice CAPTCHAs, automated phone line challenges, and strict protocol checks designed to filter out software callers. This cat-and-mouse game will shift the focus toward model alignment and trust frameworks. Developers will have to contend with complex ethical queries, such as those explored in our breakdown on aligned to whom?, as voice agents are forced to balance user fidelity against legal and platform compliance.
Monetization and the Shift Beyond Per-Token Pricing
For developers and startup founders, the commercial entry of Instinct and Metas Muse into the phone market highlights a massive transition in business models. The traditional SaaS model of charging per seat or charging pure API token markups doesn't match the economics of voice execution. Telecom minutes cost real money, carrier routing involves regional regulatory surcharges, and high-reliability low-latency voice pipelines carry significant infrastructure overhead.
Instead, we are seeing a pivot toward outcome-based and task-based pricing. Users don't want to pay for the raw minutes an agent spends sitting on hold; they want to pay a premium when the dinner reservation is secured or when the subscription is officially canceled. This aligns the incentives of the agent developer directly with the user, but it also places immense engineering pressure on reliability. If an agent fails to navigate an IVR tree after ten minutes of airtime, the platform absorbs the financial loss.
Furthermore, voice tools are quickly becoming table stakes across the application layer. Much like the audio transformations we analyzed in The Mic Is On. So Is Live Auto-Tune, low-latency live audio processing is moving from specialized hardware into ubiquitous cloud software. Founders building in the agent space need to realize that basic voice input/output is no longer a differentiator; the true defensibility lies in workflow execution, zero-trust authentication, and integration with legacy enterprise APIs.
Regulatory Minefields and the Identity Crisis in Digital Voice
We cannot talk about AI phone agents without addressing the regulatory dynamic. Telecommunication networks are among the most heavily regulated spaces on earth. Between the Telephone Consumer Protection Act
This analysis was inspired by a story originally reported by TechCrunch. Read the original report →
Supercharge Your Workflow with Claude AI
The AI assistant used by professionals worldwide. Write, code, analyse — all in one place.