LIFEHUBBER
Choose theme

AI Resources

MOSS-Audio

GitHub stars: 676 GitHub forks: 46 Last pushed September 6, 2026: Pushed 25d ago
Stats from GitHub

MOSS-Audio turns audio into text answers about speech, sounds, music and timing. The family comes from MOSI.AI, the OpenMOSS team and Shanghai Innovation Institute.

The publisher provides 4B and 8B Instruct and Thinking variants, Python examples, a Gradio interface, fine-tuning instructions and an SGLang serving guide. This is an audio-understanding family, rather than a text-to-speech generator. Use this as a first read, not a recommendation. Open the original project before trusting details like terms, limits, privacy, cost, setup, or safety.

What it is

Unified audio-understanding models

Models for asking questions about recordings, describing sounds and music, and producing transcripts with word or sentence timestamps.

Why it stands out

Questions about what happened when

The publisher describes explicit time markers between audio representations to support time-based questions and event localization alongside transcription.

Availability

Repository with model and serving paths

The repository links four checkpoints and documents local inference, an interactive app, SGLang service requests, and LoRA or full-parameter training.

Why it matters

What makes it useful

The Python infer.py example asks “Describe this audio,” even though it names the decoded text transcription. It passes both the loaded recording and that prompt to the processor before generation. Changing only AUDIO_PATH leaves the description request in place. If the task is to write down spoken words instead, the text prompt also needs to ask for transcription.

Notable points

What stands out

For a client displaying the answer, the SGLang guide shows content and reasoning_content as separate message fields. The qwen3 parser separates the think-tagged portion; the instruction-injection parser also removes the transition sentence used by thinking-budget control. The guide pairs that latter parser with its instruction-injection budget processor, so response parsing is part of the serving configuration.

Before using

What to review

The fine-tuning guide’s JSONL sample pairs a WAV path and user prompt with the assistant’s target text. Only the assistant response contributes to the training loss.

For the default PyTorch data collator, the publisher requires per_device_train_batch_size of 1 because audio spectrogram lengths vary.

For audio longer than 30 seconds, the guide directs readers to allow enough max_len; it says over-limit token sequences and spectrograms are automatically truncated.

Reader fit

Who may find it relevant

Builders adapting the acoustic encoder have an additional training choice. The fine-tuning guide lists lora_on_audio_encoder as false by default and documents an explicit flag to add LoRA to its query, key and value projections. Enabling LoRA and enabling those encoder adapters are separate settings, so an acoustic-adaptation recipe needs that second choice recorded.

Editorial note

Why LifeHubber lists it

For a live answer display, the serving guide shows stream_reasoning=false to keep reasoning tokens out of the token-by-token stream. It places answer text in delta.content and reasoning in delta.reasoning_content. That is a display choice: the guide documents thinking_budget separately for limiting thinking tokens. Choose the stream behavior and the thinking limit as two settings.

Source links

Source materials

Reader note

Before relying on this entry

LifeHubber lists entries to help readers inspect AI projects, not to endorse them or prove they are safe, suitable, accurate, maintained, or right for a specific use. We do not verify every entry in depth. Before relying on anything listed, review the original materials, terms, privacy practices, limits, and risks that matter for your situation.

Advertisements

Advertisements

For project maintainers

Listed here? You can use the badge.

If you maintain a project with a current LifeHubber listing, you may add the optional “Listed on LifeHubber AI Resources” badge to its README, docs, or website. No introduction or permission request is needed.

See what’s moving