EchoChat

EchoChat is an audio-conditioned conversational model with text and speech-token outputs. This repository provides the language model, tokenizer, and audio encoder weights.

Files and loading

  • Root: language model weights, tokenizer, configuration and custom model code.
  • audio_encoder/: complete audio encoder weights, configuration and its custom implementation.
  • LICENSE, licenses/, THIRD_PARTY_NOTICES.md: applicable licenses and notices.
import torch
from transformers import AutoModelForCausalLM, AutoTokenizer

repo = "ddwang2000/EchoChat"
tokenizer = AutoTokenizer.from_pretrained(repo, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    repo, trust_remote_code=True,
    torch_dtype=torch.bfloat16,
    attn_implementation="flash_attention_2",
).to("cuda").eval()

The custom audio encoder class is WhisperEncoder in audio_encoder/modeling_audio_encoder.py. Instantiate it with the supplied configuration and strictly load audio_encoder/model.safetensors. Its custom attention implementation requires a compatible CUDA/FlashAttention runtime.

This is a component checkpoint, not a standard Transformers text-generation pipeline. End-to-end audio use requires compatible feature extraction, prompting, and text/speech-token generation logic. Waveform synthesis additionally requires a separately licensed compatible decoder, which is not included here.

Dependencies

The component configuration targets Transformers 4.51.3. Runtime dependencies include PyTorch, Transformers, safetensors, tiktoken, einops, and a compatible FlashAttention build for CUDA execution. Inspect the custom code before enabling trust_remote_code=True.

Downloads last month
341
Safetensors
Model size
10B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support