Back to Model List

ControlFoley – Xiaomi's Open-Source Controllable Video Sound Effect Generation Model

AI Tech Editorial
RSS Feed
ControlFoley – Xiaomi's Open-Source Controllable Video Sound Effect Generation Model official screenshot
(Image source: official screenshot)

Executive Summary:

ControlFoley is an open-source controllable video sound effect generation model from Xiaomi Research, designed to solve the long-standing controllability challenge in video-to-audio (V2A). The model u...

1. What Is ControlFoley

ControlFoley is an open-source controllable video sound effect generation model from Xiaomi Research, designed to solve the long-standing controllability challenge in video-to-audio (V2A). The model uniformly supports three video dubbing tasks: text-guided dubbing (TV2A), text-controlled dubbing (TC-V2A), and reference-audio-controlled dubbing (AC-V2A). Through the proprietary spatiotemporal audio-visual encoder CAV-MAE-ST, temporal-timbre decoupling mechanism, and modality-robust training strategy, it achieves comprehensive improvements in semantic alignment, temporal synchronization, and audio quality. ControlFoley reaches open-source SOTA on multiple benchmarks including VGGSound-Test and Kling-Audio-Eval, with code and model weights fully released.

controlfoley official website screenshot
Image source: Official article

Technical positioning and domain: ControlFoley belongs to the multimodal generation field, specifically the video sound effect generation (V2A) sub-direction. Unlike traditional unconditional or text-only V2A models, ControlFoley emphasizes "controllability"—users can precisely control generated sound effect style, content, and temporal rhythm through text instructions or reference audio, offering high practical value in short video creation, film post-production, and similar scenarios.

Development background: The model was led by Xiaomi Research's YJX team. As a leading consumer electronics and AIoT company, Xiaomi has deep technical accumulation in audio-video processing and on-device inference. ControlFoley's motivation stems from existing V2A models' controllability deficiencies—most support only text guidance; when text conflicts with visual semantics, models often ignore text intent and cannot leverage reference audio for timbre control. The Xiaomi team solved these pain points through a unified framework.

Core value: ControlFoley's core value is "one model, three control modes." Creators need not switch tools—they can achieve text-prompt supplementation of missing sound semantics, forced text instruction following when text conflicts with visuals, and direct timbre style replication or adaptation via reference audio—all within the same inference framework. This flexibility greatly lowers the barrier to video sound effect production, especially for non-professional users seeking quick high-quality dubbing.

Technical characteristics: The proprietary CAV-MAE-ST spatiotemporal audio-visual encoder combined with CLIP visual encoder balances semantic understanding and temporal synchronization; temporal-timbre decoupling ensures reference audio affects only timbre without disrupting video rhythm; modality-robust training through random modality dropout and unified representation alignment enables stable output under single or multimodal input. These innovations form ControlFoley's technical moat.

2. Key Features

  • TV2A (text-guided video dubbing): Generates sound effects synchronized with on-screen actions based on input video and text prompts. Text supplements missing or ambiguous sound semantics in the footage—for example, adding "gentle rain patter" text to a silent rain scene generates rhythm-matching rain sounds. Suitable for scenes needing enhanced natural ambient sound.

  • TC-V2A (text-controlled video dubbing): When text and video semantics conflict, the model prioritizes text intent while maintaining temporal synchronization with video rhythm. For example, when a person taps a desk in video but text instructs "drum sounds," the model ignores tapping semantics and generates drum sounds at each tap action point. Solves the traditional V2A "vision-dominant" pain point where text control fails.

  • AC-V2A (reference-audio-controlled video dubbing): Users provide reference audio; the model extracts global timbre features (instrument type, timbre texture) and generates synchronized sound effects for video in that style. Temporal-timbre decoupling suppresses rhythm information from reference audio, retaining only timbre style without disrupting original action beats. Provides efficient means for brand sound unification and style transfer.

  • Unified multi-task inference framework: Single model architecture supports all three task types without switching models or fine-tuning. Internal conditional encoders distinguish input modality combinations (video only, video+text, video+reference audio), automatically adapting at inference. Reduces deployment and maintenance costs, especially for production environments needing flexible control mode switching.

  • Spatiotemporal audio-visual encoder CAV-MAE-ST: Proprietary joint visual-audio encoder focused on capturing spatiotemporal correspondence between video actions and sounds. Complements standard CLIP encoder: CLIP handles high-level semantic alignment (e.g., "cat meowing"), CAV-MAE-ST focuses on low-level temporal synchronization (precise alignment of cat mouth opening with meow). Improves audio-visual sync precision across benchmarks.

  • Modality-robust training strategy: Randomly drops partial modalities during training (e.g., randomly masking video frames or text), forcing the model to extract information from any available modality. Unified multimodal representation alignment (REPA) objective significantly improves semantic consistency across different modality combinations. Ensures reasonable sound generation even with incomplete input (video only, no text).

3. How to Use

  1. Obtain open-source resources: Visit the ControlFoley GitHub repository (https://github.com/xiaomi-research/controlfoley) to download the complete codebase and pretrained weights. The official Hugging Face model hub (https://huggingface.co/YJX-Xiaomi/ControlFoley) also provides directly loadable weight files. Try the online demo on the project website (https://yjx-research.github.io/ControlFoley_web_page/) to quickly understand effects.

  2. Environment setup and dependency installation: Recommended Python 3.8+ on Linux with NVIDIA GPU (at least 16GB VRAM, A100 or RTX 4090 recommended). Install dependencies from root requirements.txt including PyTorch 2.0+, torchaudio, transformers, opencv-python, etc. Requires decord for video decoding and vocoder components (such as HiFi-GAN). Recommend conda isolated environment to avoid version conflicts.

  3. Select task mode: Before running inference, determine task type based on creative needs. Main entry is run_inference.py, specifying mode via --task parameter: tv2a (text-guided), tc_v2a (text-controlled), or ac_v2a (reference-audio-controlled). Different modes have different input requirements—organize data accordingly.

  4. Prepare input conditions: All modes require video files (mp4, avi, etc., recommended resolution ≤1080p). TV2A and TC-V2A modes also need text prompts (via --text parameter as string or text file path). AC-V2A mode requires reference audio files (wav or mp3, 16kHz sample rate recommended). Note: TC-V2A mode requires text-visual semantic conflict to demonstrate advantage; if text matches visuals, behavior degrades to TV2A.

  5. Execute generation inference: Example command: python run_inference.py --task tv2a --video demo.mp4 --text "gentle piano music" --output_dir results. The model first samples video frames (default 8 fps), extracts spatiotemporal features via CAV-MAE-ST encoder, then generates corresponding-length audio through diffusion model. Inference time depends on video length and GPU performance—typically ~30 seconds for 10-second video on A100.

  6. Export and post-processing: Results save in --output_dir as generated .wav audio files. The model converts latent features to waveforms via VAE Decoder and Vocoder (HiFi-GAN). Recommend FFmpeg or professional video editing software to merge generated audio with original video. Generated audio length strictly aligns with video length—no manual adjustment needed.

4. Pros and Cons

Pros
Unified multi-task framework: Single model covers TV2A, TC-V2A, and AC-V2A without tool switching or fine-tuning, significantly reducing deployment and usage complexity.
Precise audio-visual sync: Proprietary CAV-MAE-ST encoder combined with CLIP leads open-source competitors in both semantic alignment and temporal sync—DeSync metric only 0.36-0.38 in conflict scenarios.
Unique reference audio control: Temporal-timbre decoupling lets reference audio affect only timbre without disrupting video rhythm—a capability unique among open-source V2A models.
Thoroughly open source with active community: Code, model weights, and technical report fully open; detailed documentation on Hugging Face and GitHub facilitates researcher reproduction and improvement.

5. Comparison with Similar Tools

Dimension ControlFoley MMAudio HunyuanVideo-Foley
Core architecture Diffusion + CAV-MAE-ST + CLIP dual encoder Diffusion + audio encoder Diffusion + video encoder
Task coverage Unified TV2A/TC-V2A/AC-V2A Mainly TV2A basic dubbing Mainly TV2A basic dubbing
Text conflict handling Strong: DeSync only 0.36-0.38, prioritizes text Weak: text easily overridden by visual info Weak: limited text control
Reference audio control Supported with temporal-timbre decoupling preserving sync Not supported Not supported
Audio-visual sync precision Excellent: CAV-MAE-ST enhances spatiotemporal correspondence, sync error <0.1s Good: audio encoder-based, ~0.2s sync error Good: video encoder-based, ~0.15s sync error
Open-source status Code, weights, technical report fully open Code and weights open Code and weights open
Community ecosystem GitHub 500+ stars, high Hugging Face downloads GitHub 200+ stars GitHub 300+ stars

Selection advice: If creative scenarios require precise sound effect style control (e.g., brand ads needing unified timbre), ControlFoley's AC-V2A mode is the only choice—other competitors don't support reference audio control. For basic short video dubbing where text matches visual semantics, MMAudio or HunyuanVideo-Foley suffice with lighter deployment. For advanced creators handling text-visual conflicts (such as animation dubbing), ControlFoley's TC-V2A mode offers highest controllability while FoleyCrafter performs moderately. Overall, ControlFoley leads comprehensively in controllability and sync precision—suitable for professional users with high audio-visual quality requirements.

6. Editor's Take

ControlFoley excels in technical innovation, especially the combination of CAV-MAE-ST spatiotemporal encoder and temporal-timbre decoupling—solving two long-standing V2A pain points: control when text and visual semantics conflict, and video rhythm disruption when introducing reference audio. Compared to MMAudio and HunyuanVideo-Foley supporting only basic TV2A, ControlFoley unifies three control modes—a first among open-source models. From practical value, the model directly lowers professional barriers to video sound effect production—creators need no audio editing skills, generating synchronized sound effects through text or reference audio alone, especially suitable for short video and content creator high-frequency production. However, high hardware requirements (A100 recommended) limit adoption, and Chinese support needs community optimization. Target users include AI content creators, film post-production professionals, game sound designers, and multimodal researchers. Strong future potential: with on-device inference optimization and lightweight versions, integration into phones and smart devices for real-time sound generation is possible. Recommended rating: ★★★★☆ (4.5/5)—deductions for hardware barrier and Chinese support, but overall technical leadership and open-source completeness make it one of the best current V2A choices.

7. Use Cases

  • Short video creation: Creators add customized sound effects to silent footage via text prompts (e.g., "upbeat electronic music") generating rhythm-matching audio. To unify timbre across a video series, reference audio control mode quickly replicates previous video sound effect texture without manual adjustment each time.

  • Animation and game sound design: Generate specific-style impact or ambient sounds for character actions. For example, animation shows character knocking on wooden door but text instructs "metal impact sound"—TC-V2A mode ignores visual semantics, generating metal sounds at each knock point. Game developers can use reference audio to replace ordinary footsteps with robot mechanical sounds.

  • Film post-production dubbing: In movie or series post-production, unify sound effect style across scenes. AC-V2A mode applies reference audio (such as specific-era phonograph timbre) to multiple scenes maintaining sound consistency. CAV-MAE-ST encoder ensures audio-visual sync precision, reducing manual adjustment workload.

  • Advertising and marketing: Brand ads quickly generate rhythm-synchronized dubbing matching brand tone from text instructions. Input "luxurious, deep cello sound" and the model generates corresponding sound effects following on-screen rhythm, strengthening emotional expression. Reference audio control replicates competitor or own classic ad timbre for style transfer.

  • Creator live stream clip editing: After live streams, creators convert highlight clips to short videos but original audio may contain noise. TV2A mode generates clean sound effects based on visuals (cheering, applause) improving content polish. Modality-robust training ensures reasonable background sound generation even with video-only input.

8. FAQ

Q: Does ControlFoley have video resolution requirements?
A: Recommended resolution ≤1080p. The model internally scales frames (default 256×256 processing)—higher resolution increases decode time without improving generation quality. For 4K video, downsample before input.

Q: Is generated audio duration exactly matching video?
A: Yes. CAV-MAE-ST encoder obtains video frame timestamps, then generates corresponding-length audio. Fixed 16kHz sample rate, duration precisely aligned with video—no post-stretching needed.

Q: In TC-V2A mode, when text conflicts with visuals, how does the model maintain temporal sync?
A: CAV-MAE-ST encoder extracts action rhythm from video (knock frequency, motion amplitude) encoded as temporal conditions. Text instructions affect only timbre and semantic content; temporal conditions force audio generation aligned with action points. So even with semantic conflict, sound effects appear at correct action moments.

Q: Can I use Chinese text prompts? How effective is it?
A: Yes, but less effective than English. Training data primarily from English datasets (such as VGGSound)—Chinese semantic understanding is limited. Recommend simple Chinese descriptions (e.g., "rain sound"); complex instructions may be misunderstood. Community is advancing Chinese fine-tuned versions.

Q: How do I merge generated audio with video?
A: Recommend FFmpeg: ffmpeg -i video.mp4 -i generated.wav -c:v copy -c:a aac -map 0:v:0 -map 1:a:0 output.mp4. Also usable in Adobe Premiere, CapCut, and other video editing software by directly importing audio tracks.

Q: How fast is model inference? Can it be used for live streaming?
A: On A100 GPU, ~30 seconds for 10-second video. Current version cannot meet real-time requirements (<1 second)—better for offline production. Future lightweight versions or TensorRT optimization may enable real-time inference.

9. Project Links

Related AI Model Articles

© All Rights Reserved. Some content on this site is partially generated by AI with human review.