Atlas TTS by runAtlas on Pipeshift: Multilingual speech on your own GPUs

Production-grade multilingual neural TTS you can run on your own GPUs, live on Pipeshift in days.

CEO at Pipeshift

Pipeshift

Author

Published

Topic

Partnership

Atlas TTS is now available on Pipeshift. The 1.7B-parameter neural TTS model exposes real-time streaming, preset and custom voices, and an OpenAI-compatible API.

Pipeshift gives teams two ways to run the same model: a managed endpoint or a deployment inside their own VPC. Teams can validate voice quality and integration first, then choose the operating model that fits their traffic and data requirements.

Production TTS becomes an infrastructure decision once usage grows. Live voice traffic has to stay responsive under concurrent calls, while regulated deployments may need synthesis text and audio to remain inside a controlled boundary. Those requirements determine GPU capacity and the deployment model that fits the workload.

What does Atlas TTS on Pipeshift include?

Atlas TTS on Pipeshift combines Atlas’s model and voice layer with Pipeshift’s deployment and inference infrastructure.

  • runAtlas provides the 1.7B-parameter TTS model, preset and custom voices, voice cloning, pronunciation tuning, and enterprise support.

  • Pipeshift provides the serving layer around the model, including GPU deployment, endpoint delivery, scaling, and managed or customer-VPC infrastructure.

The application can keep the same OpenAI-compatible request shape across both deployment paths. Teams can validate voice quality and integration on a managed endpoint, then move the same client integration into their own environment when the workload requires it.

What is Atlas TTS?

Atlas TTS is runAtlas’s 1.7B-parameter neural text-to-speech engine for real-time and long-form speech. It supports sentence-by-sentence streaming and can run as a managed endpoint or inside customer-controlled infrastructure.

Which languages and voice options does Atlas TTS support?

One engine supports English, German, French, Russian, Portuguese, Spanish, Italian, Chinese, Japanese, and Korean, plus dialects. Atlas keeps the same voice identity across languages, so a product can use one branded voice across markets instead of maintaining separate language-specific voice stacks.

Atlas provides named preset voices, zero-shot voice cloning from about three seconds of reference audio, and custom brand voices. Enterprise deployments can also include pronunciation tuning and dedicated support from the Atlas team.

How does Atlas TTS integrate with existing applications?

Atlas streams audio sentence by sentence over WebSocket and returns WAV or PCM at 24 kHz, 16-bit mono. The API follows the OpenAI speech interface, with official OpenAI SDK support for Python and Node, so an existing client can keep the same request pattern and point to the Atlas endpoint.

How can you deploy Atlas TTS on Pipeshift?

Atlas TTS keeps the same request shape across managed and self-hosted deployments. In practice, the application can stay stable even when the serving boundary changes.

The operational difference is who owns the inference infrastructure and how capacity is paid for. The VPC path also keeps synthesis text, voice references, and generated audio inside the customer environment at inference time.

When should you use the managed endpoint?

Use the managed endpoint when the priority is getting a pilot live without owning the inference stack. Pipeshift handles hosting, scaling, and billing behind the OpenAI-compatible endpoint, which also fits traffic that changes sharply over time.

When should you run Atlas TTS in your VPC?

Use the VPC path when synthesis data needs to stay inside the customer environment or when sustained volume makes customer-owned capacity preferable. Atlas supports deployments with no callbacks, telemetry, or phone-home at inference time, including air-gapped environments. The VPC track uses an annual license and customer-controlled GPU infrastructure rather than per-character metering.

The same request shape is available on both paths, so a team can start managed and move the application into its own VPC without rewriting the client integration.

How does Atlas TTS perform on H100 and B300 GPUs?

Atlas publishes performance at rated concurrency, so the benchmark reflects the conditions a production voice service has to sustain rather than an idle single-stream test. Put another way, the useful question is how many streams one GPU can sustain while holding the latency target.

GPU

Rated load

Speed / stream

TTFB p50

Aggregate

H100

8 streams

3.5× RT

115 ms

~30× real-time

B300

10 to 12 streams

4 to 5× RT

79 ms

~46× real-time

How to read the H100 and B300 benchmark?

The 38 ms single-stream result measures best-case time to first audio for one request. It is useful for understanding model responsiveness, but it is not the number to use when sizing a production service.

For capacity planning, the rated-load results are more useful. The H100 sustains eight streams at 115 ms p50 TTFB and about 30× aggregate real-time. The B300 sustains 10 to 12 streams at 79 ms p50 TTFB and about 46× aggregate real-time. In Atlas’s published sweep, the B300 supports more concurrent streams while maintaining lower p50 latency.

How does self-hosting change TTS costs?

A VPC deployment uses an annual software license with customer-controlled GPU infrastructure instead of per-character metering. The cost question moves from characters generated to the amount of serving capacity the workload needs. For a team already reserving GPUs, that could look like another serving budget for the infrastructure team sizes and monitors.

What replaces per-character pricing in a self-hosted deployment?

The run-rate is the software license plus the GPU time required to serve the workload. Higher synthesis volume consumes more of that GPU capacity, so cost scales through infrastructure rather than a separate character meter.

How should you estimate GPU capacity?

Start from expected peak concurrency and the rated streams each GPU can sustain at the target latency. Then account for utilization and operational headroom.

  • Peak concurrency sets the highest number of simultaneous synthesis streams the deployment needs to serve.

  • Per-GPU stream capacity comes from the rated concurrency each GPU can sustain at the latency target.

  • Utilization determines how much of the reserved GPU time turns into useful synthesis work over the day.

  • Headroom covers traffic spikes and leaves room for operational variance without pushing the service beyond its rated load.

The H100 and B300 results above provide the per-GPU starting point. A team can map expected concurrency to GPU count, estimate GPU-hours, and compare that run rate with the managed endpoint using the same workload assumptions.

How should you evaluate Atlas TTS on your own scripts?

Voice quality should be tested on the content the application will actually speak. That could be something as simple as the scripts that already expose awkward pronunciation or pacing in the current voice.

You can test Atlas TTS instantly in the Atlas TTS playground using your own scripts.

Atlas offers a blind comparison using customer scripts, with vendor identities removed and the customer’s own team scoring the outputs.

Start with a blind comparison against the incumbent voice

  1. Send representative production scripts, including difficult names, domain terms, long-form passages, and multilingual content where relevant.

  2. Render Atlas TTS and the incumbent provider side by side with vendor identities removed.

  3. Have the same internal reviewers score naturalness, pacing, pronunciation, and voice consistency. Atlas returns the comparison within two working days.

Separate voice quality from runtime performance

The blind comparison answers whether the voice holds up on the content you ship. A production pilot should then measure the serving system under the concurrency the application expects.

  • Time to first audio and streaming stability show whether the voice stays responsive during live interactions.

  • Sustained throughput and GPU utilization show how much traffic the deployment can serve before additional capacity is required.

Application metrics such as handle time, contact rate, abandonment, or cost per call can then be evaluated on a real slice of traffic rather than inferred from a synthetic benchmark.

Which workloads are a good fit for Atlas TTS?

Atlas TTS is a fit when latency, voice consistency, or deployment control affects the product experience. The same engine can cover live and long-form speech without forcing teams to maintain separate language-specific serving stacks.

  • Live voice agents and IVR systems can stream speech sentence by sentence while teams measure first-audio latency at the concurrency callers will actually create.

  • Contact centers and high-volume outbound workflows can use the same serving stack for conversational speech where cost per contact changes what can run at full scale.

  • Regulated voice workloads in BFSI, healthcare, and telecom can keep synthesis text, voice references, and generated audio inside customer-controlled infrastructure.

  • Localization and dubbing workflows can keep one voice identity across the ten supported languages plus dialects.

  • Accessibility, narration, and embedded products can use the same OpenAI-compatible integration for long-form or in-product speech.

How do you move from evaluation to production on Pipeshift?

Teams can start with a blind comparison on their own scripts, then validate latency and integration on the traffic pattern they expect. The managed endpoint is the shortest route to a pilot because Pipeshift handles hosting and scaling behind the OpenAI-compatible endpoint.

If the workload needs customer-controlled infrastructure, the same application integration can move into the VPC path without a new client. Pipeshift handles deployment and inference infrastructure, while Atlas provides the model, voices, pronunciation tuning, and enterprise support.

Talk to Pipeshift about the deployment path and GPU capacity that fit your workload.

Frequently asked questions

Can Atlas TTS run on-prem or in an air-gapped environment?

Atlas TTS can run inside customer-controlled infrastructure, including VPC, on-prem, and air-gapped environments. That keeps synthesis text, reference voices, and generated audio inside the customer boundary at inference time.

For isolated deployments, Atlas supports operation without callbacks, telemetry, or phone-home requirements. Pipeshift provides the deployment and serving infrastructure around the model, while the customer retains control of the environment where inference runs.

Can I use Atlas TTS with the OpenAI SDK?

Yes. Atlas follows the OpenAI speech interface and supports the official OpenAI SDKs for Python and Node. An existing application can keep the same request pattern and point the client to the Atlas endpoint.

Can I move from a managed Atlas TTS endpoint to my own VPC without changing the client integration?

Yes. The managed and self-hosted paths use the same request shape, so the application does not need a separate TTS integration when the deployment boundary changes.

The infrastructure behind that request does change. A managed endpoint puts hosting and scaling behind Pipeshift, while the VPC path moves inference into customer-controlled infrastructure and makes GPU capacity part of the customer deployment.

How is the self-hosted Atlas TTS priced?

Self-hosted Atlas TTS uses an annual software license with customer-controlled GPU infrastructure rather than per-character metering. The serving run-rate therefore depends on the GPU capacity reserved for the workload.

Teams can estimate that capacity from peak concurrency, the rated number of streams each GPU can sustain, utilization, and operational headroom. The H100 and B300 benchmark results provide a starting point for that calculation.

Can Atlas TTS use a custom brand voice?

Yes. Atlas supports named preset voices, zero-shot voice cloning from about three seconds of reference audio, and custom brand voices. Enterprise deployments can also include pronunciation tuning for terms that need consistent handling in production.

Talk to Pipeshift about deploying Atlas TTS by runAtlas

Talk to Pipeshift about deploying Atlas TTS by runAtlas

Talk to Pipeshift about deploying Atlas TTS by runAtlas

By Pipeshift

©2026 Infercloud Inc.

By Pipeshift

©2026 Infercloud Inc.