wilsonix-studio
其他 活跃维护

wilsonix-studio

ewceniza9009/wilsonix-studio

AI技术的桌面数字音频工作站,内置音轨分离、和弦检测、卡拉OK功能,无需复杂配置即可在本地完成音频编辑、人声提取、伴奏生成等操作,轻量化部署适配个人创作者日常使用

0
Stars 标星
0
Forks 分支
0
Watchers 关注
0
Open Issues
Other
主要语言
None
开源协议
16 KB
仓库大小
17 天前
最后推送
一键安装扩展 / 插件指令
dsh plugin --profile web add github:ewceniza9009/wilsonix-studio
git clone https://github.com/ewceniza9009/wilsonix-studio.git
git clone git@github.com:ewceniza9009/wilsonix-studio.git
README.md main
# WilSonix Studio PRO **AI-Powered Desktop DAW — Stem Separation, Karaoke & Live Performance Suite** ![Version](https://cdnimage-cache.doubi.ren/?url=https://img.shields.io/badge/version-0.2.4-blue) ![Platform](https://cdnimage-cache.doubi.ren/?url=https://img.shields.io/badge/platform-Windows%20%7C%20macOS%20%7C%20Linux%20%7C%20Android-brightgreen) ![License](https://cdnimage-cache.doubi.ren/?url=https://img.shields.io/badge/license-Proprietary-red) *Separate. Align. Master. Perform.*

Download

Windows InstallerDownload v0.2.4 from Releases

Platform Status
Windows x64 (NSIS Installer) Stable (v0.2.4)
Android (Capacitor APK) Beta
macOS / Linux Build from source

What's New in v0.2.4

  • Stem-Targeted Harmonies: Generate 3-part diatonic harmonies on either LEAD_VOCALS or BACKGROUND_VOCALS directly from the mixer channel strip.
  • Dynamic Stem Indicator: Vocal Studio HUD displays the exact active target stem and musical key.
  • Go Gateway Proxy Fix: Direct reverse-proxy forwarding through port 7860 for all harmony, spectral repair, and vocal character endpoints.
  • Continuous Phase Vocoder Auto-Tune: Zero audio slicing, zero chunk boundaries, zero boundary clicks, and zero comb-filtering.
  • High-Speed 3-Part Vocal Harmonizer: Generates +3rd Diatonic, +5th Power, and -8va Sub-Vocal stems in seconds.
  • Adaptive Light/Dark Mode: Vocal Studio and Spectral Repair modals dynamically restyle for both light and dark themes.
  • Consonant & Sibilance De-Ess Guard: Prevents sizzling/sparkling phase fuzz by automatically detecting and bypassing pitch-shifting on unvoiced consonants (s, sh, t, k, p, breaths), keeping them 100% pure acoustic audio.
  • AI Singer Gender & Duet Detection: Real-time vocal acoustic classifier detects Male, Female, and Duet/Polyphonic vocal intervals.
  • 3-Tier Color-Coded Karaoke Teleprompter: Words dynamically light up in Electric Neon Cyan (Male), Radiant Hot Pink (Female), and Golden Amber (Duet / Both), accompanied by active singer badges on the TV stage prompter ([♂ SINGER 1], [♀ SINGER 2], [👥 DUET / BOTH]).
  • Interactive Singer Override: Clickable [♂ HE] / [♀ SHE] / [👥 DUET] pills on every lyric line in the teleprompter editor allow instant manual reassignment with automatic saving.
  • 1080p MP4 Video Produce with Burned Multi-Style Subtitles: Advanced SubStation Alpha (.ass) rendering engine burns exact singer colors directly into exported videos via FFmpeg.
  • Continuous Pitch Performance Scoring: Overhauled scoring engine evaluates pitch accuracy and rhythm on 50ms hop windows across the full vocal take with cents-level vibrato tolerance (no hardcoded scores).
  • AI Spectral Repair & Stem "Healing Brush": Interactive 2D STFT spectrogram viewer ($20\text{ Hz} - 20\text{ kHz}$) with click-and-drag marquee box selection and 3 surgical inpainting algorithms: Attenuate (-18dB), Ambient Fill (Wiener), and Harmonic Interpolation.
  • Next-Gen Vocal Studio & Diatonic 3-Part Harmonizer: Generates discrete +3rd Diatonic, +5th Power, and -8va Sub-Vocal stems locked to the detected song key and chords, adding them directly into the DAW mixer as controllable faders.
  • Formant-Locked Voice Character Morphing: Apply Female Bright, Male Warmth, or Vintage Radio profiles without chipmunk distortion.
  • Native Rust DSP Core (crates/wilsonix-dsp): SIMD RealFFT inpainting engine and GCC-PHAT transient phase alignment for zero-latency C-speed audio computation.

Features

Stem Separation

Upload any audio or video file and let AI split it into isolated stems:

Engine Stems Speed Best For
Demucs v4 6 (Vocals, Drums, Bass, Guitar, Piano, Other) ~2.5 min Full studio decomposition
Demucs v4 4 (Vocals, Drums, Bass, Other) ~1.5 min Classic multi-track
UVR MDX-Net ONNX 2 (Vocal + Backing) ~50 sec Fast karaoke / minus-one
BS-RoFormer 2 (Vocal + Backing) SOTA (12.98 dB SDR) Studio mastering quality
  • 32-bit float processing throughout
  • Multi-guitar spatial decomposition (4-way azimuth)
  • Solo Protect — recovers lead instruments from vocal stem
  • Quality modes: Draft / Studio HQ / Ultra HD / Mastering Grade

DAW Mixer

Full-featured multi-channel mixer with per-stem controls:

  • Volume Fader (0–150%) + Pan Knob (stereo panner)
  • Mute / Solo per channel
  • 3-Band Parametric EQ — Low shelf (250 Hz), Mid peaking (1.2 kHz), High shelf (4.5 kHz)
  • Low-Cut Filter — 80 Hz rumble removal
  • Compressor — Threshold -18 dB, ratio 3:1
  • Tape Saturator — Analog tube warmth, 4x oversampling
  • Stereo Chorus — LFO-modulated delay
  • Haas Effect Widener — Stereo spatial expansion
  • Reverb Send — Convolution hall bus (2.2s)
  • Delay Send — 350ms tempo-synced feedback delay
  • Per-channel FFT analyser + OLED waveform display
  • 6 built-in FX presets per channel
  • Import custom WAV files as mixer channels

Chord Detection

  • Real-time chord recognition powered by AI
  • Instrument selector: Guitar / Bass / Piano / All (Mix)
  • Scrollable chord progression timeline in the Harmonics strip
  • Chords cached in database for instant reload

Karaoke Stage

  • Whisper AI word-level transcription (millisecond precision)
  • Multi-language: English, Bisaya, Tagalog, Japanese, Korean, Chinese, Spanish, French, German
  • OLED teleprompter with physics bouncing ball + word-by-word glow highlight
  • 6 trance visual themes (Cyber Aurora, Warp Tunnel, Plasma Orbs, Synthwave Grid, Matrix Starfield, Kaleidoscope Nebula)
  • Lyrics editor with undo/redo, gap auto-fill, hallucination detection panel
  • Export: .LRC and .ASS subtitle files
  • Sync offset nudge (-0.1s / +0.1s)
  • Live microphone sing-along with real-time pitch tracking
  • Pitch radar canvas visualization

Auto-Tune & Vocal Processing

  • Neural Auto-Tune (WORLD Vocoder) — formant-preserving pitch correction
  • Scales: Chromatic, Auto, Major keys
  • Presets: T-Pain, Pop, Natural
  • AI Mic Enhancement — Denoise + De-Ess + Capsule Exciter
  • Profiles: Neumann U87, Shure

Dual-Screen TV Stage

  • Pop-out dedicated fullscreen teleprompter to any external TV or monitor
  • Zero-latency lyrics teleprompter with chord sync
  • Perfect for live karaoke performance

Karaoke Video Production

  • 1080p MP4 export with burned-in animated glowing lyrics
  • Highlight colors: Neon Cyan, Warm Amber, Vibrant Pink, Emerald Mint
  • Vocal guide volume slider
  • Title / Artist metadata
  • Performance video remux (browser recording → broadcast H.264/AAC)

Auto-Mastering

  • One-click AI mastering to -14 LUFS (streaming standard)
  • ITU-R BS.1770 / EBU R128 K-weighting filter
  • Multi-band DSP processing
  • True-Peak limiter
  • 24-bit mastered WAV output

File Inspector

  • MIDI file parser — notes, tempo map, piano roll visualization
  • ASS subtitle inspector — styles, events, karaoke timings
  • Quick ASS override tag parser

Backend Logs

  • Real-time log streaming with color-coded output
  • Filter by level: All / Errors / Warnings / Info
  • Auto-scroll with clear button

Library & Project History

  • Search and filter past projects
  • Status tracking (completed / failed / processing)
  • One-click "Open in Studio" to reload stems
  • Download stems as ZIP
  • Auto-refresh polling for active jobs

Admin Control Center

  • CPU, RAM, GPU, Disk telemetry
  • User management (roles: admin / producer)
  • Job queue monitoring
  • Stem purge for expired data

Tech Stack

Layer Technology Purpose
Desktop Shell Tauri v2 (Rust) Native window, system integration, process management
Backend Server Go HTTP server, WebSocket relay, SQLite DB, auth, routing
AI Engine Python 3.12 + PyTorch Stem separation, chord detection, lyrics transcription
ML Models Demucs v4, UVR MDX-Net, BS-RoFormer, OpenAI Whisper AI audio processing
Frontend Vanilla JS + Tailwind CSS + Anime.js + Lucide Icons Reactive DAW workspace
Audio DSP Web Audio API (32-bit float) Real-time mixing, EQ, compression, effects
Database SQLite (via Go) Users, jobs, stems, lyrics, chords
Mobile Capacitor + ONNX Runtime Mobile Android APK with offline inference
Installer NSIS Windows installer packaging
Build PyInstaller Python → single .exe bundling

System Requirements

Component Minimum Recommended
OS Windows 10 x64 Windows 11
RAM 4 GB 8 GB+
Storage 2 GB free 5 GB+ (for models)
CPU Dual-core 2 GHz Quad-core+ with AVX
GPU None (CPU works) NVIDIA GPU with CUDA for faster inference

Supported Formats

Input Formats
Audio WAV, FLAC, MP3, M4A, OGG, OPUS
Video MP4, MKV, MOV, WebM, AVI
Subtitles .ASS, .SSA
MIDI .mid, .midi

Keyboard Shortcuts

Key Action
Space Play / Pause
M Mute selected channel
S Solo selected channel
0 Rewind to beginning
[ / ] Set Loop Point A / B
L Toggle Loop
Ctrl+Z Undo
Ctrl+Shift+Z Redo
Esc Close modal

Mobile (Android)

WilSonix Studio PRO is available as an Android APK via Capacitor:

  • Swipeable 6-stem mixer carousel
  • Gesture-locked audio sliders
  • Offline AI inference with ONNX Runtime Mobile
  • Touch-optimized UI with safe-area support

Building from Source

Prerequisites

Build Steps

# Clone
git clone https://github.com/ewceniza9009/pygo.git
cd pygo

# Python environment
python -m venv python_env
python_env\Scripts\activate
pip install -r requirements.txt

# Build Python worker (PyInstaller)
pyinstaller --onefile --name python_worker --noconfirm \
  --distpath src-tauri\resources \
  --collect-data audio_separator \
  --collect-submodules audio_separator \
  --collect-all onnxruntime \
  run_worker.py

# Build Go server
go build -ldflags "-s -w" -o src-tauri\resources\pygo_server.exe .

# Build Tauri desktop app
cd src-tauri
cargo tauri build

Architecture

┌─────────────────────────────────────────────┐
│           Tauri (Rust) Shell                │
│  ┌──────────┐  ┌──────────────────────────┐ │
│  │ Webview  │  │   Process Manager        │ │
│  │ (HTML)   │  │   ├─ pygo_server.exe     │ │
│  │          │  │   └─ python_worker.exe   │ │
│  └──────────┘  └──────────────────────────┘ │
└─────────────────────────────────────────────┘
         │ HTTP/WebSocket        │ HTTP
    ┌────▼────┐            ┌────▼────┐
    │   Go    │────────────│ Python  │
    │ Server  │  proxy WS  │ FastAPI │
    │ :7860   │            │ :19876  │
    └────┬────┘            └────┬────┘
         │                      │
    ┌────▼────┐            ┌────▼────────────┐
    │ SQLite  │            │ PyTorch / ONNX  │
    │  (DB)   │            │ Demucs/Roformer │
    └─────────┘            │ Whisper         │
                           └─────────────────┘

Developer

Erwin Wilson E. Ceniza


License

Proprietary. All rights reserved.