On September 10, 2026, Phony.ai announced a voice-agent platform built on an unusual promise: the customer picks the carrier, the language model, the speech engine and the voice, and none of those choices are welded to the platform itself. That announcement, covered in the Phony.ai coverage on manilatimes.net, is one of the clearer signs the voice-agent category is entering an unbundling phase.
The trouble is that "provider-neutral" has become marketing wallpaper. Almost every vendor claims some version of it, and buyers are picking up a lot of half-truths along the way. Here are six of the most common ones, and what's going on underneath.
Myth: Provider-Neutral Just Means "We Support Multiple LLMs"
The label gets slapped on any platform that lists a few model names on a dropdown. Real neutrality means the platform doesn't depend on any specific model or provider for its core behavior, and can integrate, evaluate and switch between them without an application rewrite. Inbenta describes this as a layered design where the customer-facing channels and the governed knowledge stay stable while the model layer is treated as an interchangeable component.
Apply that test to voice and the question gets sharper. A dropdown that lets you pick between two hosted LLMs isn't neutrality if the speech-to-text engine, the voice, the turn-taking logic and the telephony are all still owned by the same vendor. Neutrality has to reach every layer that affects a call, or it's a marketing word.
Myth: Unbundling Is a Downgrade From an Integrated Stack
The pitch for an all-in-one stack is that everything is tuned together, so latency and quality beat anything you could bolt together yourself. There's truth in that, right up until the business hits a ceiling it can't move. A Deepgram guide on unbundling makes the point plainly: bundled platforms tend to run into pricing and compliance walls when provider choice or data-plane control starts to matter, and events like a compliance audit, a multi-region latency requirement or a bring-your-own-model mandate are what force the move.
Both architectures work. They just fail in different places. Unbundling is a downgrade only if you never hit the ceiling.
Myth: The Model Is the Only Layer Worth Swapping
Model choice gets almost all the airtime. But a voice call is a pipeline, and any one link in it can be the one holding you back. The layers that get chosen and re-chosen tend to look like this:
- Speech-to-text. Accuracy on phone-quality audio, handling of accents and cross-talk, and the latency of partial transcripts drive how natural the agent feels before it says a word.
- The language model. Reasoning quality gets the headlines. Cost per token and response latency decide whether you can afford to run it on every call.
- Text-to-speech and voice. Time-to-first-audio is the single most audible piece of the stack, and it's a moving target as providers ship new models.
- The carrier. Per-minute rates, number availability and compliance posture vary widely, and none of it is technically bound to who runs the AI.
- Orchestration. Turn-taking, barge-in, handoff logic and event logs sit above the other four and determine how the whole call behaves.
Buyers who only argue about which LLM is best are optimizing one link in a five-link chain. The wins often sit somewhere else.
Myth: The Carrier Is Just a Pipe
Carrier choice is the layer most often smuggled into a bundled price and marked up without you seeing it. A platform that hides per-minute telephony inside a single blended rate is charging you for opacity.
The other half of the argument is portability. If the AI platform owns the number, changing platforms means changing numbers, and for most businesses that's not a change they'll voluntarily make. A neutral design lets the business keep its carrier relationship and its numbers, and lets the AI ride on top.
Myth: Neutrality Is About Price, Not Governance
Cost gets used as the headline argument for unbundling. It isn't the deepest one. When each layer is a separate, replaceable component, you can also see what each layer did on a given call, who processed the audio, which model generated which response, and what evidence exists that the call was permitted in the first place.
Consent rules for automated calls keep shifting, disclosure expectations for AI voices are hardening, and buyers are being asked to prove what happened on a specific call months after it ended. A stack you can inspect layer by layer is a stack you can defend.
What Buyers Should Actually Ask
If you're evaluating a voice-agent platform, the useful questions are less about which providers are on the list and more about how the layers are bound together:
- Per-layer cost. Does each call return an itemized breakdown of speech, model, voice and telephony, or a blended per-minute number?
- Bring-your-own carrier. Can you keep your existing numbers and carrier account, or does the platform insist on owning that relationship?
- Model swap path. If a better or cheaper model ships next quarter, is moving to it a config change or a project?
- Consent evidence. For every outbound call the platform places, what artifact proves the call was permitted, and where is it stored?
- Data boundary. Is your material used to instruct the agent, or to train someone's model?
Vendors that answer these plainly are the ones building for the unbundled era. The ones that hedge are hoping you don't ask again.

