This directory contains the neural audio generation pipeline for NAW.
The neural engine implements a two-stage hybrid pipeline:
┌─────────────────────────────────────────────────────────────┐
│ STAGE 1: Semantic Planner (Autoregressive Transformer) │
│ ───────────────────────────────────────────────────────── │
│ Input: Text Prompt + BPM + Control Signals │
│ Output: Coarse Musical Skeleton (Structure, Rhythm, Pitch) │
│ Speed: Fast (~2 seconds for 32 bars) │
└─────────────────────────────────────────────────────────────┘
↓
┌─────────────────────────────────────────────────────────────┐
│ STAGE 2: Acoustic Renderer (Flow Matching / Diffusion) │
│ ───────────────────────────────────────────────────────── │
│ Input: Semantic Skeleton + Text Prompt │
│ Output: High-Fidelity Audio with Realistic Timbre │
│ Speed: Slower (~10 seconds for 32 bars) │
└─────────────────────────────────────────────────────────────┘
neural-engine/
├── index.ts # Main exports
├── codec/
│ └── DACCodec.ts # DAC audio codec
├── planner/
│ └── SemanticPlanner.ts # Autoregressive planner
├── renderer/
│ └── AcousticRenderer.ts # Flow matching renderer
├── vocoder/
│ └── Vocoder.ts # Audio reconstruction (Vocos/DisCoder/HiFiGAN)
├── control/
│ └── ControlNet.ts # Fine-grained control adapters
├── conditioning/
│ └── CLAP.ts # Audio-text conditioning
└── inpainting/
└── SpectrogramInpainter.ts # Surgical audio editing
import {
DACCodec,
SemanticPlanner,
AcousticRenderer,
Vocoder,
VocoderType,
generateMusic,
} from './neural-engine';
// Simple generation (recommended)
const stems = await generateMusic({
text: "Uplifting house track",
bpm: 128,
bars: 32,
quality: 'balanced', // 'fast' | 'balanced' | 'high'
});
// Advanced usage with components
const codec = new DACCodec();
const planner = new SemanticPlanner();
const renderer = new AcousticRenderer();
const vocoder = new Vocoder({ type: VocoderType.VOCOS });
await codec.initialize();
await planner.initialize();
await renderer.initialize();
await vocoder.initialize();
// Stage 1: Generate semantic skeleton
const skeleton = await planner.generate({
text: "Energetic drum pattern",
bpm: 140,
bars: 16,
});
// Stage 2: Render acoustic details
const latents = await renderer.render({
text: "Analog warm synth",
semanticTokens: skeleton.tokens,
});
// Stage 3: Decode to audio with vocoder
const result = await vocoder.decode(latents[0]);
console.log(`Decoded ${result.audio.length} samples at ${result.rtf}x realtime`);- DAC encoder implementation (stub)
- DAC decoder implementation (stub)
- RVQ quantization (stub)
- Latent space visualization (stub)
- EnCodec compatibility layer
- Working test suite
- Transformer-XL architecture (stub)
- Multi-stream prediction (stub)
- Training data pipeline
- ONNX export and INT8 quantization
- Alternative: Mamba-based planner
- Working test suite
- Flow Matching architecture (stub)
- DiT backbone implementation (stub)
- CLAP text conditioning (stub)
- Classifier-free guidance (stub)
- Vocoder integration (Vocos/DisCoder/HiFiGAN) - Architecture complete
- TensorRT optimization
- Working test suite
- Vocoder module (Vocos/DisCoder/HiFiGAN)
- generateMusic() end-to-end pipeline
- Multi-vocoder support with quality presets
- Stem mixing and normalization
- Real-time factor reporting
- Complete test coverage
- ControlNet adapters for fine-grained control
- CLAP audio-text conditioning
- Spectrogram inpainting
- Style adapters (LoRA-based)
- Working test suite
npm test # Basic tests (53 tests)
npm run test:integration # Integration tests (25 tests)The test suite validates:
- Component initialization
- Configuration management
- Encoding/decoding pipeline
- Semantic planning
- Acoustic rendering
- Vocoder decoding and mixing
- ControlNet, CLAP, and Inpainting
- End-to-end pipeline with generateMusic()
npm run demoThe demo shows:
- Full three-stage pipeline in action
- Semantic Planning → Acoustic Rendering → Vocoding
- Progress reporting
- Stem generation with vocoder
- Full pipeline using generateMusic()
- Real-time factor measurements
All tests currently pass with stub implementations. Real neural models will be integrated progressively.
- DAC: Descript Audio Codec
- EnCodec: High Fidelity Neural Audio Compression
- MusicGen: Simple and Controllable Music Generation
- Transformer-XL: Attentive Language Models Beyond a Fixed-Length Context
- Flow Matching: Flow Matching for Generative Modeling
- Stable Audio: Fast Timing-Conditioned Latent Audio Diffusion
Compresses and decompresses audio to/from latent space using Residual Vector Quantization (RVQ).
class DACCodec {
constructor(config?: Partial<DACConfig>)
// Initialize the codec (loads model weights)
async initialize(): Promise<void>
// Encode audio to latent tokens
async encode(audio: Float32Array): Promise<DACLatent>
// Decode latent tokens back to audio
async decode(latent: DACLatent): Promise<Float32Array>
// Extract semantic tokens (first 2 codebooks)
extractSemanticTokens(latent: DACLatent): number[][]
// Extract acoustic tokens (remaining codebooks)
extractAcousticTokens(latent: DACLatent): number[][]
}Configuration Options:
interface DACConfig {
sampleRate: number; // Audio sample rate (default: 44100 Hz)
latentRate: number; // Latent representation rate (default: 24000 Hz)
numCodebooks: number; // Number of RVQ codebooks (default: 16)
codebookSize: number; // Size of each codebook (default: 1024)
semanticCodebooks: number; // Semantic codebooks count (default: 2)
}Generates coarse musical structure using an autoregressive transformer.
class SemanticPlanner {
constructor(config?: Partial<SemanticPlannerConfig>)
// Initialize the planner (loads model weights)
async initialize(): Promise<void>
// Generate semantic skeleton from prompt
async generate(
prompt: SemanticPrompt,
onProgress?: (progress: number) => void
): Promise<SemanticSkeleton>
// Set temperature for randomness control
setTemperature(temperature: number): void
}Configuration Options:
interface SemanticPlannerConfig {
modelSize: 'small' | 'medium' | 'large';
contextWindow: number; // Token context length (default: 2048)
temperature: number; // Sampling temperature (default: 0.8)
topK: number; // Top-K sampling (default: 50)
topP: number; // Nucleus sampling (default: 0.95)
useKVCache: boolean; // Enable KV-cache for speed (default: true)
}Renders high-fidelity audio from semantic tokens using flow matching.
class AcousticRenderer {
constructor(config?: Partial<AcousticRendererConfig>)
// Initialize the renderer (loads model weights)
async initialize(): Promise<void>
// Render audio from semantic tokens
async render(
prompt: RenderPrompt,
onProgress?: (progress: RenderProgress) => void
): Promise<DACLatent[]>
// Set guidance scale for text conditioning
setGuidanceScale(scale: number): void
// Set quality preset
setQualityPreset(preset: 'fast' | 'balanced' | 'high'): void
}Configuration Options:
interface AcousticRendererConfig {
qualityPreset: 'fast' | 'balanced' | 'high';
numSteps: number; // Diffusion steps (default: 20)
guidanceScale: number; // CFG guidance strength (default: 3.0)
modelSize: 'base' | 'large'; // Model size
useVocoder: boolean; // Use integrated vocoder (default: true)
vocoderType: VocoderType; // Vocoder backend
}Converts latent representations to audio waveforms.
class Vocoder {
constructor(config?: Partial<VocoderConfig>)
// Initialize the vocoder (loads model weights)
async initialize(): Promise<void>
// Decode latent to audio
async decode(latent: DACLatent): Promise<VocoderResult>
// Decode and mix multiple stems
async decodeMix(
latents: DACLatent[],
volumes?: number[]
): Promise<VocoderResult>
}Vocoder Types:
VocoderType.VOCOS: Fast preview (25x realtime)VocoderType.DISCODER: High-quality final renderVocoderType.HIFIGAN: Alternative for compatibility
Fine-grained control over generation using control signals.
class ControlNet {
constructor(config?: Partial<ControlNetConfig>)
async initialize(): Promise<void>
// Extract control signal from audio
async extractControlSignal(
audio: Float32Array,
type: ControlType
): Promise<ControlSignal>
// Apply control to latent
async applyControl(
latent: number[][],
signal: ControlSignal,
strength?: number
): Promise<number[][]>
// Load style adapter (LoRA)
async loadStyleAdapter(name: string): Promise<StyleAdapter>
}Control Types:
ControlType.MELODY: Melodic contour controlControlType.RHYTHM: Rhythmic pattern controlControlType.DYNAMICS: Loudness envelope controlControlType.TIMBRE: Spectral characteristicsControlType.HARMONY: Harmonic progression
Audio-text contrastive conditioning for reference-based generation.
class CLAP {
constructor(config?: Partial<CLAPConfig>)
async initialize(): Promise<void>
// Encode audio to embedding
async encodeAudio(audio: Float32Array): Promise<AudioEmbedding>
// Encode text to embedding
async encodeText(text: string): Promise<TextEmbedding>
// Compute similarity between audio and text
computeSimilarity(
audioEmbed: AudioEmbedding,
textEmbed: TextEmbedding
): number
// Blend audio and text embeddings
blendEmbeddings(
audioEmbed: AudioEmbedding,
textEmbed: TextEmbedding,
audioWeight: number
): Float32Array
}Surgical audio editing through inpainting.
class SpectrogramInpainter {
constructor(config?: Partial<InpaintingConfig>)
async initialize(): Promise<void>
// Inpaint masked region
async inpaint(
audio: Float32Array,
mask: InpaintingMask
): Promise<InpaintingResult>
// Extend audio (outpainting)
async outpaint(
audio: Float32Array,
extendSeconds: number,
seamlessLoop?: boolean
): Promise<Float32Array>
// Detect loop points
async detectLoopPoints(
audio: Float32Array
): Promise<Array<{ start: number; end: number; confidence: number }>>
}import { generateMusic, loadStyleAdapter } from './neural-engine';
// Generate 4 stems with different style adapters per stem
const stems = await generateMusic({
text: "Electronic music",
bpm: 128,
bars: 32,
quality: 'balanced',
stemStyles: {
DRUMS: await loadStyleAdapter('techno'),
BASS: await loadStyleAdapter('techno'),
VOCALS: await loadStyleAdapter('jazz'),
OTHER: await loadStyleAdapter('orchestral')
}
});import { CLAP, generateMusic } from './neural-engine';
const clap = new CLAP();
await clap.initialize();
// Load reference audio
const referenceAudio = await loadAudioFile('reference.wav');
const audioEmbed = await clap.encodeAudio(referenceAudio);
// Generate with audio reference
const stems = await generateMusic({
text: "Similar vibe but faster tempo",
bpm: 140,
bars: 32,
audioReference: audioEmbed,
audioReferenceWeight: 0.6 // 60% audio, 40% text
});import { SpectrogramInpainter } from './neural-engine';
const inpainter = new SpectrogramInpainter();
await inpainter.initialize();
// Load audio
const audio = await loadAudioFile('track.wav');
// Define mask for region to regenerate (e.g., remove snare)
const mask: InpaintingMask = {
startTime: 2.0, // Start at 2 seconds
endTime: 2.5, // End at 2.5 seconds
freqMin: 200, // 200 Hz
freqMax: 8000, // 8000 Hz (snare frequency range)
};
// Inpaint the masked region
const result = await inpainter.inpaint(audio, mask);
console.log(`Inpainted ${result.inpaintedSamples} samples`);import { ControlNet, ControlType, AcousticRenderer } from './neural-engine';
const controlNet = new ControlNet();
await controlNet.initialize();
const renderer = new AcousticRenderer();
await renderer.initialize();
// Extract melody from reference
const referenceMelody = await loadAudioFile('melody_reference.wav');
const melodySignal = await controlNet.extractControlSignal(
referenceMelody,
ControlType.MELODY
);
// Generate with melody control
const latents = await renderer.render({
text: "Synthwave track",
controlSignals: [melodySignal],
controlStrength: 0.8
});import { SemanticPlanner, AcousticRenderer, Vocoder } from './neural-engine';
const planner = new SemanticPlanner();
const renderer = new AcousticRenderer();
const vocoder = new Vocoder();
await Promise.all([
planner.initialize(),
renderer.initialize(),
vocoder.initialize()
]);
// Stage 1: Semantic planning with progress
console.log('Stage 1: Semantic Planning...');
const skeleton = await planner.generate(
{ text: "Epic orchestral", bpm: 90, bars: 64 },
(progress) => console.log(`Planning: ${Math.round(progress * 100)}%`)
);
// Stage 2: Acoustic rendering with progress
console.log('Stage 2: Acoustic Rendering...');
const latents = await renderer.render(
{ text: "Epic orchestral", semanticTokens: skeleton.tokens },
(progress) => console.log(`Rendering: ${Math.round(progress.percentage)}%`)
);
// Stage 3: Vocoding
console.log('Stage 3: Vocoding...');
const audioResults = await Promise.all(
latents.map(latent => vocoder.decode(latent))
);
console.log(`Generated ${audioResults.length} stems`);
audioResults.forEach((result, i) => {
console.log(`Stem ${i}: ${result.rtf}x realtime`);
});import { SpectrogramInpainter } from './neural-engine';
const inpainter = new SpectrogramInpainter();
await inpainter.initialize();
// Load 8-bar loop
const loop = await loadAudioFile('8bar_loop.wav');
// Extend to 32 bars with seamless looping
const extended = await inpainter.outpaint(
loop,
24, // Extend by 24 bars (3x the original)
true // Seamless loop
);
console.log(`Extended from 8 to 32 bars`);- DAC Codec: ~500MB VRAM (encoder + decoder)
- Semantic Planner: ~1.5GB VRAM (500M params)
- Acoustic Renderer: ~4GB VRAM (1B params)
- Total: ~6GB VRAM minimum for full pipeline
| Component | Target Latency | Achieved (RTX 3090) |
|---|---|---|
| DAC Encode | 10ms | TBD |
| Semantic Plan (1 bar) | 30ms | TBD |
| Acoustic Render (1 bar) | 50ms | TBD |
| Vocoding | 10ms | TBD |
| Total (1 bar) | 100ms | TBD |
- Use KV-Cache: Enable in SemanticPlanner for 3-5x speedup
- Batch Processing: Generate multiple stems in parallel
- Quality Presets: Use 'fast' preset for iteration, 'high' for final render
- GPU Selection: RTX 3090 or better recommended for real-time
- Quantization: INT8 quantization provides 4x memory reduction with minimal quality loss
| Use Case | Minimum | Recommended | Optimal |
|---|---|---|---|
| Preview | GTX 1660 (6GB) | RTX 3060 (12GB) | RTX 3090 (24GB) |
| Production | RTX 3060 (12GB) | RTX 3090 (24GB) | RTX 4090 (24GB) |
| Real-time | RTX 3090 (24GB) | RTX 4090 (24GB) | A100 (40GB) |
Issue: "Model not found" error on initialization
// Solution: Ensure model weights are downloaded
await downloadModels(); // Utility to download pre-trained weightsIssue: Out of memory during rendering
// Solution: Use lower quality preset or smaller batch size
const renderer = new AcousticRenderer({
qualityPreset: 'fast', // Reduces VRAM usage
modelSize: 'base' // Use smaller model
});Issue: Slow generation on CPU
// Solution: Enable GPU acceleration
const vocoder = new Vocoder({
useGPU: true,
type: VocoderType.VOCOS // Fastest vocoder
});Issue: Clicks/pops in generated audio
// Solution: Use seamless loop mode or adjust fade settings
const inpainter = new SpectrogramInpainter({
blendingMethod: 'smooth',
fadeLength: 0.1 // 100ms fade
});Enable verbose logging to diagnose issues:
import { setDebugMode } from './neural-engine';
setDebugMode(true); // Enables detailed logging
// Now all operations will log detailed information
const stems = await generateMusic({ text: "Test", bpm: 120, bars: 8 });Currently, this directory contains stub implementations with the correct interfaces and architecture. Actual neural models will be integrated in Phase 2.
To add new features:
- Follow the existing interface patterns
- Add comprehensive JSDoc comments
- Update this README
- Add TODO comments for future implementation
- Write tests before implementing real models
- Document performance characteristics
See CONTRIBUTING.md for general guidelines. For neural-engine specific:
- Model Integration: Follow the stub→implementation→optimization pattern
- Testing: Add both unit tests and integration tests
- Documentation: Update API reference and usage examples
- Benchmarking: Include performance measurements on standard hardware
- ROADMAP.md - Development timeline
- ARCHITECTURE.md - Technical details
- CONTRIBUTING.md - Contribution guide
- COMMERCIAL_LICENSE.md - Licensing information
- Issues: GitHub Issues
- Discussions: GitHub Discussions
For understanding the underlying algorithms and architectures, see the Research References section above and ARCHITECTURE.md.