Choose theme
AI Resources
MOSS-Audio
MOSS-Audio turns audio into text answers about speech, sounds, music and timing. The family comes from MOSI.AI, the OpenMOSS team and Shanghai Innovation Institute.
The publisher provides 4B and 8B Instruct and Thinking variants, Python examples, a Gradio interface, fine-tuning instructions and an SGLang serving guide. This is an audio-understanding family, rather than a text-to-speech generator. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.
What it is
Unified audio-understanding models
Models for asking questions about recordings, describing sounds and music, and producing transcripts with word or sentence timestamps.
Why it stands out
Questions about what happened when
The publisher describes explicit time markers between audio representations to support time-based questions and event localization alongside transcription.
Availability
Repository with model and serving paths
The repository links four checkpoints and documents local inference, an interactive app, SGLang service requests, and LoRA or full-parameter training.
Why it matters
What makes it useful
The Python infer.py example asks “Describe this audio,” even though it names the decoded text transcription. It passes both the loaded recording and that prompt to the processor before generation. Changing only AUDIO_PATH leaves the description request in place. If the task is to write down spoken words instead, the text prompt also needs to ask for transcription.
What to know
Where it fits
The SGLang guide distinguishes two thinking controls. Its enable_thinking chat-template switch applies only to pure text requests; the guide says audio takes a different template branch, so that switch does not affect it. The documented thinking-budget method instead uses a custom logit processor with the corresponding server mode. Changing the text switch alone is not the documented way to limit audio thinking.
Notable points
What stands out
For a client displaying the answer, the SGLang guide shows content and reasoning_content as separate message fields. The qwen3 parser separates the think-tagged portion; the instruction-injection parser also removes the transition sentence used by thinking-budget control. The guide pairs that latter parser with its instruction-injection budget processor, so response parsing is part of the serving configuration.
Before using
What to review
The fine-tuning guide’s JSONL sample pairs a WAV path and user prompt with the assistant’s target text. Only the assistant response contributes to the training loss.
For the default PyTorch data collator, the publisher requires per_device_train_batch_size of 1 because audio spectrogram lengths vary.
For audio longer than 30 seconds, the guide directs readers to allow enough max_len; it says over-limit token sequences and spectrograms are automatically truncated.
Reader fit
Who may find it relevant
Builders adapting the acoustic encoder have an additional training choice. The fine-tuning guide lists lora_on_audio_encoder as false by default and documents an explicit flag to add LoRA to its query, key and value projections. Enabling LoRA and enabling those encoder adapters are separate settings, so an acoustic-adaptation recipe needs that second choice recorded.
Editorial note
Why LifeHubber lists it
For a live answer display, the serving guide shows stream_reasoning=false to keep reasoning tokens out of the token-by-token stream. It places answer text in delta.content and reasoning in delta.reasoning_content. That is a display choice: the guide documents thinking_budget separately for limiting thinking tokens. Choose the stream behavior and the thinking limit as two settings.
Source links
Source materials
Reader note
Before relying on this entry
LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.
More in Speech Models
Keep browsing this category
Explore more speech model resources.
Fish Audio S2 Pro
fishaudio/s2-pro
A text-to-speech model with detailed control over prosody and emotional delivery.
AuK
Tencent-Hunyuan/AuK
A 1.5B speech model for instruction-guided text-to-speech, content and acoustic editing, paralinguistic changes, speech enhancement, and source separation, with public code, weights, demos, ComfyUI nodes, and fine-tuning materials.
LFM JP
LiquidAI/lfm-jp
A Liquid AI Hugging Face collection for Japanese-tuned LFMs, grouping LFM2.5-1.2B-JP-202606 and LFM2.5-Audio-1.5B-JP materials for Japanese text, tool use, structured outputs, ASR, TTS, and speech-to-speech workflows.
For project maintainers
Listed here? You can use the badge.
If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.