Files

339 lines
7.3 KiB
Markdown

# Qwen3-TTS OpenAI-like API
A FastAPI-based OpenAI-compatible API for Qwen3 Text-to-Speech with streaming support and voice cloning capabilities.
## Features
- 🎵 **Real-time Streaming** - Stream audio chunks as they're generated
- 🎭 **Voice Cloning** - Clone voices using reference audio
- 🌍 **Multi-language Support** - English, Russian, Chinese, Japanese, Korean, and more
- 📡 **OpenAI-like API** - Familiar API structure for easy integration
- 🚀 **High Performance** - Optimized streaming with CUDA support
- 📚 **Interactive Docs** - Swagger UI at `/docs`
## Installation
### Prerequisites
- Python >= 3.9
- CUDA 11.8+ (for GPU support) or CPU mode
- 8GB+ VRAM recommended for GPU
### Setup
1. **Install dependencies**
```bash
pip install fastapi uvicorn gradio requests soundfile librosa numpy torch transformers
```
2. **Install Qwen3-TTS package**
```bash
pip install -e .
```
## Quick Start
### Starting the API Server
```bash
# Run the API server
python -m api.main
# Or with uvicorn directly
uvicorn api.main:app --host 0.0.0.0 --port 8000
```
The API will be available at `http://localhost:8000`
### Running the Gradio Demo
In a separate terminal:
```bash
python -m api.gradio_demo
```
The demo will be available at `http://localhost:7860`
## API Endpoints
### Health Check
```
GET /v1/health
```
Check API status and model readiness.
**Response:**
```json
{
"status": "healthy",
"model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
"device": "cuda",
"ready": true
}
```
### List Models
```
GET /v1/models
```
Get available TTS models.
**Response:**
```json
{
"object": "list",
"data": [
{
"id": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
"object": "model",
"owned_by": "Alibaba",
"supported_languages": ["English", "Russian", "Chinese", "Japanese", "Korean"],
"supports_streaming": true,
"supports_voice_clone": true
}
]
}
```
### Text-to-Speech (Standard)
```
POST /v1/audio/speech
```
Generate audio from text.
**Request:**
```json
{
"text": "Hello, how are you?",
"language": "English",
"model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
"voice_clone_mode": "disabled",
"speed": 1.0,
"pitch": 1.0
}
```
**Response:**
```json
{
"id": "req_abc123def456",
"object": "audio",
"created": 1708284000,
"model": "Qwen/Qwen3-TTS-12Hz-1.7B-Base",
"audio_base64": "UklGRi...",
"duration": 2.5,
"sample_rate": 24000,
"language": "English"
}
```
### Text-to-Speech (Streaming)
```
POST /v1/audio/speech/stream
```
Stream audio chunks as they're generated.
**Request:**
```json
{
"text": "Hello, this is a streaming test.",
"language": "English",
"voice_clone_mode": "disabled",
"stream_options": {
"emit_every_frames": 8,
"decode_window_frames": 80,
"overlap_samples": 512
}
}
```
**Response:**
- Binary stream of WAV frame chunks
- Each chunk is prefixed with a 4-byte length (big-endian)
- Format: `[length (4 bytes)][WAV frame data]...`
## Voice Cloning
### Setup
Voice cloning requires two files in `assets/voice_cloning/`:
1. **`ref_audio.wav`** - Reference audio file
- Format: WAV
- Duration: 5-30 seconds
- Clear, natural speech
2. **`ref_text.txt`** - Transcription of reference audio
- Plain text UTF-8
- Must exactly match the audio content
### Using Voice Cloning
```bash
# 1. Place reference files
cp your_reference.wav assets/voice_cloning/ref_audio.wav
echo "Reference text here..." > assets/voice_cloning/ref_text.txt
# 2. Start the API (it will load voice cloning automatically)
python -m api.main
# 3. Use voice cloning in requests
curl -X POST "http://localhost:8000/v1/audio/speech" \
-H "Content-Type: application/json" \
-d '{
"text": "New text to synthesize",
"language": "English",
"voice_clone_mode": "reference_audio"
}'
```
## Python Integration
### Using the API from Python
```python
import requests
import base64
from io import BytesIO
import soundfile as sf
# Generate speech
response = requests.post(
"http://localhost:8000/v1/audio/speech",
json={
"text": "Hello, world!",
"language": "English",
"voice_clone_mode": "disabled"
}
)
data = response.json()
audio_bytes = base64.b64decode(data["audio_base64"])
audio, sr = sf.read(BytesIO(audio_bytes))
# Save to file
sf.write("output.wav", audio, sr)
```
### Streaming Example
```python
import requests
response = requests.post(
"http://localhost:8000/v1/audio/speech/stream",
json={
"text": "This is a streaming test.",
"language": "English",
},
stream=True
)
chunks = []
for chunk in response.iter_content(chunk_size=65536):
# Parse frame length (4 bytes)
frame_len = int.from_bytes(chunk[:4], byteorder='big')
# Extract WAV frame data
frame_data = chunk[4:4 + frame_len]
chunks.append(frame_data)
```
## Configuration
### Model Selection
Currently supports:
- `Qwen/Qwen3-TTS-12Hz-1.7B-Base` (default)
More models can be added to `TTSModel` enum in `api/models.py`.
### Device Selection
The API automatically selects the best available device:
- CUDA GPU (if available and with sufficient VRAM)
- CPU (fallback)
To force a specific device, modify `api/main.py`:
```python
# In ModelState.__init__
self.device = "cpu" # Force CPU mode
```
### Streaming Options
Adjust streaming performance in `api/models.py`:
```python
class StreamOptions(BaseModel):
emit_every_frames: int = 8 # Lower = more chunks, lower latency
decode_window_frames: int = 80 # Higher = better quality, higher latency
overlap_samples: int = 512 # Overlap for smooth transitions
```
## Performance Tips
1. **GPU Optimization**
- Use CUDA 11.8+ for better performance
- `flash_attention_2` is automatically enabled when available
2. **Streaming Optimization**
- Lower `emit_every_frames` for faster first chunk
- Increase `decode_window_frames` for better audio quality
3. **Production Deployment**
- Use a production ASGI server like Gunicorn or Hypercorn
- Enable response caching for identical requests
- Monitor GPU memory usage
## Troubleshooting
### API won't start
```
error: No module named 'transformers'
→ pip install transformers accelerate
```
### Model loading fails
```
error: CUDA out of memory
→ Use CPU mode: modify device to "cpu" in api/main.py
→ Or reduce batch size
```
### Voice cloning not working
- Verify `ref_audio.wav` and `ref_text.txt` exist in `assets/voice_cloning/`
- Check file permissions
- Ensure reference text matches audio content exactly
### Slow responses
- Check GPU/CPU utilization
- Reduce `decode_window_frames` for faster (but lower quality) generation
- Ensure sufficient system memory available
## API Documentation
Once the server is running, visit:
- **Interactive Docs (Swagger UI)**: `http://localhost:8000/docs`
- **ReDoc Documentation**: `http://localhost:8000/redoc`
## Examples
See the `examples/` directory for more usage examples:
- `test_streaming.py` - Streaming generation example
- `test_model_12hz_base.py` - Base model usage
- `test_model_12hz_custom_voice.py` - Voice cloning example
## License
This API is licensed under the Apache 2.0 License. See LICENSE file for details.
## Contributing
Contributions are welcome! Please feel free to submit pull requests.
## Support
For issues, questions, or suggestions:
1. Check the troubleshooting section
2. Review the API documentation at `/docs`
3. Check logs for detailed error messages