Skip to content

Latest commit

 

History

History
495 lines (365 loc) · 24.9 KB

File metadata and controls

495 lines (365 loc) · 24.9 KB

Plugins

See also: Architecture

A plugin registers ops, stats, or methods on audio.

// my-plugin.js
import audio from 'audio'

const myOp = (input, output, ctx) => {
  for (let c = 0; c < input.length; c++)
    for (let i = 0; i < input[c].length; i++)
      output[c][i] = -input[c][i]
}

audio.op('myOp', myOp)                       // shorthand for { process: myOp }
// or explicitly:
audio.op('myOp', { process: myOp })
import 'my-plugin.js'   // registers op + wires a.myOp()

Ops

A block processor: (input, output, ctx) => void.

input and output are separate Float32Array[] per channel, BLOCK_SIZE samples (1024 default, last block may be shorter). Read from input, write to output — never assume they alias. The engine pre-allocates two buffer sets and rotates them per op in the pipeline chain (previous output becomes next input). Zero allocation in the hot path.

ctx persists across chunks — set any property for stateful computation. Fixed fields update per chunk: at, duration, sampleRate, blockOffset, totalDuration, plus named params and any extras from the edit object.

When an op declares params, positional arguments are mapped to named properties on ctx. Prefer params so processors read explicit names (ctx.value, ctx.freq, etc.) rather than positional arrays.

  • at/duration — op time range (seconds), chunk-relative
  • totalDuration — full audio duration (seconds)
  • blockOffset — absolute position of this chunk (seconds)
  • channel — which channel(s) the op is scoped to

Convert to samples:

let sr = ctx.sampleRate
let start = ctx.at != null ? Math.round(ctx.at * sr) : 0
let end = ctx.duration != null ? start + Math.round(ctx.duration * sr) : input[0].length

For passthrough (no-op), copy input to output:

for (let c = 0; c < input.length; c++) output[c].set(input[c])

Options

By default, edits are stored as ['myOp', opts].

Any plain object as the last argument is treated as options. Known keys (at, duration, channel) are extracted; sample-based aliases offset/length convert to at/duration. All other keys flow through as extras and arrive in ctx:

a.fade(1, { curve: 'exp' })
// → edit: ['fade', { in: 1, curve: 'exp' }]
// → ctx.curve === 'exp'

With params, named params can live in the options object too — no positional args needed:

a.gain({value: -6, at: 0.5})
// → edit: ['gain', { value: -6, at: 0.5 }]
// → ctx.value === -6, ctx.at === 0.5

Querying ops

audio.op('gain')  // → descriptor { process, ... } or undefined
audio.op()        // → all ops: { gain: {...}, crop: {...}, ... }

Descriptor

Each op is a descriptor object with stage handlers and options. Pass a function for the shorthand process-only form, or an object for the full form:

audio.op('myOp', myProcess)                          // shorthand for { process: myProcess }
audio.op('myOp', { process, plan, resolve, ... })    // full descriptor

Stage handlers (each op defines one or more):

audio.op('myOp', {
  params: ['arg1', 'arg2'],               // named positional arguments → ctx.arg1, ctx.arg2
  process: (input, output, ctx) => { },   // per-block PCM transform (read input, write output)
  plan: (segs, ctx) => segs,              // structural segment rewrite
  expand: (ctx) => edit,                  // macro: rewrite into simpler edit(s), no stats
  resolve: (ctx) => edit,                 // stat-conditioned: replace using decoded stats
  ranged: true,                           // op handles {at, duration} itself — engine skips its range scoping
  auto: 'sample',                         // op samples function params itself (default: engine, 128-sample steps)
  fnArgs: ['arg1'],                       // params that are genuine functions, not automation (e.g. transform's fn)
  prepare: async (a, index) => { },       // async work before render (model inference), see below
})

process ops get {at, duration} range scoping, function-param automation, and click-free value ramps from the engine automatically — declare ranged/auto/fnArgs only to opt out or take over.

params

Declare named positional arguments. The first positional arg maps to ctx[params[0]], the second to ctx[params[1]], etc. This applies to all stages: process, plan, and resolve.

audio.op('gain', {
  params: ['value'],
  process: (input, output, ctx) => {
    let g = 10 ** (ctx.value / 20)
    for (let c = 0; c < input.length; c++)
      for (let i = 0; i < input[c].length; i++)
        output[c][i] = input[c][i] * g
  }
})

With params, calling with a single options object works naturally — named params and range opts coexist:

a.gain({value: -6, at: 0.5})  // named param + range opt → ctx.value = -6
a.eq({freq: 1000, Q: 2})      // multiple named params

Positional args override opts for the same param. If both a.gain(-6, {value: -3}) are present, the positional -6 wins.

plan

Rewrite the segment map without touching PCM. For ops that change timeline geometry (crop, insert, remove, repeat, pad, reverse, speed).

compilePlan(a, len, final) compiles a.edits into a segment map + sample pipeline + limit. Edits are the source of truth; segments are the compiled form — like bytecode from source, or DOM patches from VDOM. Segments are never maintained manually — they rebuild on any edit change.

During streaming (final=false), compilePlan is called repeatedly as more source data arrives. Each call recompiles all edits from scratch and tracks a limit — the safe output boundary given current source length. adjustLimit(limit, type, ctx) transforms the limit per op. When final=true (fully decoded), limit equals totalLen.

ctx has total, sampleRate, offset, length, plus named params from params. The offset/length are at/duration pre-converted to samples (null if unset).

import { seg } from 'audio/plan.js'

audio.op('myRepeat', { plan(segs, ctx) {
  let r = [...segs]
  for (let s of segs) { let n = s.slice(); n[2] = s[2] + ctx.total; r.push(n) }
  return r
} })

Most structural ops already have reusable segment transforms you can import instead of writing raw segment math:

import { cropSegs } from 'audio/fn/crop.js'       // cropSegs(segs, offset, length)
import { insertSegs } from 'audio/fn/insert.js'   // insertSegs(segs, at, length, ref)
import { removeSegs } from 'audio/fn/remove.js'   // removeSegs(segs, offset, duration)
import { reverseSegs } from 'audio/fn/reverse.js' // reverseSegs(segs, offset, end)
import { speedSegs } from 'audio/fn/speed.js'     // speedSegs(segs, rate)

Segment format

A segment is a copy instruction: [from, count, to, rate?, ref?, interp?].

Read count samples from source at from, write to output at to. All offsets are absolute — segments are independent, not linked. You can process them in any order, binary search by output position, or skip segments for partial renders.

Index Field Description
0 from Read offset in source (samples)
1 count Number of samples to copy
2 to Write offset in output (samples)
3 rate Source read rate. Omit or 1 = forward. -1 = reverse (used by reverse()). 2 = read 2× faster, halving duration (used by speed()). 0.5 = half speed, doubling duration. The speed op multiplies existing rates and adjusts count — speed(2) on a 10s segment produces count/2 at rate*2. The renderer uses linear interpolation for non-unit rates unless interp is set.
4 ref Source: undefined = self, null = zero-fill (silence), audio instance = external
5 interp Optional interpolation function for non-unit rate reads. It receives (src, target, tOff, n, rate, phase) and may expose .margin for source context.
6 env Optional fade envelope: flat pairs [p0, p1, …] of fade phase at the segment's start and end, linear between; gain Π sin(p·π/2). Envelope segments are summed, not written, so a fade-out and fade-in over the same span form an equal-power crossfade (remove/insert with crossfade).

seg(from, count, to, rate?, ref?, interp?, env?) creates a segment; subSeg(s, at, n, to) cuts one to a sub-range (envelope included); derive segments through it rather than by hand.

Examples

10s audio at 44100 Hz starts as one segment — the whole source maps 1:1 to output:

[0, 441000, 0]       →  read all 441000 samples from 0, write at 0

After crop({at: 2, duration: 3}) — keep only 3s starting at the 2s mark:

[88200, 132300, 0]   →  read 132300 samples from 88200, write at 0

After insert(silence, {at: 1}) — split at 1s, insert 1s silence, shift the rest:

[0, 44100, 0]              →  first 1s unchanged
[0, 44100, 44100, , null]  →  1s silence (ref=null means zero-fill)
[44100, 396900, 88200]     →  remainder shifted right by 1s

After reverse() — same range, negative rate:

[0, 441000, 0, -1]  →  read backwards

expand

Pure macro expansion — rewrite this op into simpler edits, no stats needed. Same return contract as resolve below, same ctx minus stats. Use it whenever the rewrite depends only on parameters (fade in+out → two fades, stretch → segment rate + DSP stage, resample → _resample_seg, crossfade → pad + blend). Reading expand vs resolve in a descriptor tells you at a glance whether an op needs decoded audio before it can plan.

audio.op('fade', {
  params: ['in', 'out'],
  process: fade,
  expand: (ctx) => typeof ctx.out === 'number'
    ? [['fade', { in: ctx.in }], ['fade', { in: -Math.abs(ctx.out) }]]
    : null  // single-sided — fall through to own process
})

whole

Whole-render processing — same (input, output, ctx) signature as process, called once with the entire signal (streaming: false contract plugins register this way via audio.use). At compile the engine materializes the plan so far, runs the hook, and continues from the result as a reference segment — downstream ops apply to the processed output, undo unwinds the single edit, and the materialization is cached per version. During a live decode the safe limit stays 0 (the whole signal isn't known yet); output begins once decode completes. Bounded by the flat-render guard (2³⁰ samples). Cannot be emitted by expand/resolve. A streaming: false plugin given { at, duration } reads the whole signal and keeps its output in the range alone; the hook gets ctx.at/ctx.duration to do so itself.

sidechain key

A contract plugin declaring more than one input bus (channels: { inputs: [2, 2] }) receives bus 1 from the op's key option — an audio instance or Float32Array[], rendered per block at the op's timeline position, sample-rate-reconciled, page-primed and awaited like any ref:

music.ducker({ key: voice, threshold: -30 })

load

A lazy module for the op, resolved before any plan compiles. The atom stays out of the bundle until an edit uses it; process/whole read it from desc.mod:

audio.op('repair', {
  params: ['band'],
  load: () => import('@audio/denoise-repair'),
  process: (input, output, ctx) => { let repair = audio.op('repair').mod.default /* … */ }
})

prepare

async (a, index) => void, awaited per edit after load, before anything renders: work an edit needs that cannot run inside the synchronous render, such as model inference over its whole input. The edit's input is a as the edits before it leave it; the op keeps its result on the edit's options under a symbol key, which reaches ctx and stays out of serialization. vocals({ model }) separates this way, then its process copies the result block by block. A prepare that needs only part of its input reads that range, and a source still arriving is waited for only that far: denoise({ noise: { at, duration } }) learns its noise print from the range, then streams (its input audio.from(a), which follows a source still arriving, with a's edits before index).

latency

Declared lookahead in samples — a number, or (opts, sampleRate) => samples for param-dependent delay lines. The engine compensates at the plan level (plugin delay compensation): render cursors run the pipeline's total latency ahead of the requested timeline, so delayed output lands aligned and the final samples flush through past-the-end silence. Contract plugins declare latency per CONTRACT.md and get this automatically.

warmup

Input the op must see before an output position to produce it: a number of samples or (opts, sampleRate) => samples. Seeks and ranged reads start the pipeline that much earlier and discard the lead-in, so a read from the middle equals the same span of a full render. An STFT op warms up over the frames overlapping its first output; a repair over the context it interpolates from.

holdback

How much output depends on the unknown end of a live stream: (opts, sampleRate, total) => samples, for pipeline and resolve ops alike (total for a decision that needs all of it). While the source is still arriving, output stops that far before the current end and resumes as more arrives, so a week-long stream with a fade-out holds back the fade's length, not the week. Ranges counted from the end ({at: -2}) hold back by themselves; declare holdback for end-relative parameters the engine can't see:

audio.op('fade', { holdback: (o, sr) => o.in < 0 && o.at == null ? -o.in * sr : 0, /* … */ })

streaming

A live source (audio() + push, a pipe, a socket, a response body) renders as it arrives, and its output equals the whole-file render. Every op streams, holding back only its declared latency, warmup and holdback, except whole ops, plugins registered streaming: false, reverse() of everything and the clipboard: those wait for the end, and so does a decision that reads the whole input (declared as a holdback of all of it: trim's automatic threshold, normalize's one gain for the selection). A resolve op after an op its stats can't be derived through (trim after a filter) gets the stats of the stage before it, accumulated as that stage renders. Nothing is processed past what is settled, the latency pre-roll included. test/stream.js checks each chain live against the whole file and keeps the list of ops that still wait; that list only shrinks.

resolve

Pre-render replacement using decoded stats (stat-conditioned — trim, normalize). The engine remaps stats through any structural edits already in the chain, so ctx.stats is always in this op's own output space. During incremental streaming, ctx.stats may come from stats.snapshot() with partial: true — resolve can return partial results that refine as more data decodes (e.g. trim detects head silence early, normalize applies gain from available peaks). Output stops at the stats horizon, and a decision the rest of the stream can still change is held back: trim holds a silent tail until sound resumes or the stream ends, shrink holds an open pause at its target gap.

audio.op('trim', {
  params: ['threshold'],
  process: trim,
  resolve: (ctx) => {
    let { stats, sampleRate, totalDuration, threshold } = ctx
    if (!stats?.min) return null  // no stats — fall back to per-page
    // ...analyze stats to find silence boundaries...
    return ['crop', { at: start / sampleRate, duration: (end - start) / sampleRate }]
  }
})

ctx has stats, sampleRate, channelCount, channel, at, duration, totalDuration, plus named params and edit extras. Return:

  • edit(s) — replace this op with simpler op(s)
  • false — skip (no change needed)
  • null — fall back to per-page processing

resolve runs at render time with decoded audio stats and replaces abstract ops with concrete ones.

Once the whole signal is known (ctx.final), ctx.measure(edits) renders the plan so far plus candidate pipeline edits and returns their block stats: a what-if pass for decisions stats alone can't make exactly. normalize uses it to land loudness through its true-peak limiter, after stepping on a block model; its adaptive mode, which decides while a stream is live, stays on the model. An op that measures holds its output until the end (holdback), so nothing it decides is emitted before it measures.

pointwise

Mark an op as a pure, monotonic per-sample transform: output depends only on input value, not position or history.

audio.op('clamp', {
  pointwise: true,
  params: ['limit'],
  process: (input, output, ctx) => {
    let limit = ctx.limit
    for (let c = 0; c < input.length; c++)
      for (let i = 0; i < input[c].length; i++)
        output[c][i] = Math.max(-limit, Math.min(limit, input[c][i]))
  }
})

The engine auto-derives min/max/clipping stats by probing process with edge values — no full stream recompute needed after edits. a.stat('db') resolves instantly. Energy, mean square and DC of a nonlinear map aren't functions of block extremes, so queries that need them (loudness, rms, dc) render instead.

Don't use for stateful ops (filters) or position-dependent ops (fades, automation).

For advanced cases where rms/dc/energy need algebraic precision, use deriveStats: (stats, opts) => {} instead — see gain and dc ops for examples.

Persistent ctx

ctx is the same object across all chunks — any property you set persists. Fixed fields (at, blockOffset) update each chunk; everything else stays. This handles algorithmic state like IIR filter memory:

const filter = (input, output, ctx) => {
  if (!ctx.z) ctx.z = input.map(() => 0)  // init once, persists across chunks
  for (let c = 0; c < input.length; c++) {
    output[c].set(input[c])
    // ...use ctx.z[c] for filter memory, mutate output[c] in-place
  }
}

audio.op('filter', { process: filter })

When seeking mid-stream, the engine silently renders 8 prior blocks to warm up stateful ops before producing output.

Stats

Register a stat descriptor:

audio.stat('mystat', {
  block: (chs, ctx) => chs.map(ch => /* number */),
  reduce: (blockValues, from, to) => { let v = 0; for (let i = from; i < to; i++) v += blockValues[i]; return v },
})

Or shorthand (block-only, no scalar/binned query):

audio.stat('mystat', (chs, ctx) => chs.map(ch => /* number */))

block is called per 1024-sample block during decode. Return number (all channels) or array (per-channel). Stored in a.stats.mystat as Float32Array[].

One pass can serve several stats: register the same block function for each and return a record keyed by stat name. The engine calls a shared function once per block; the built-in min/max/dc/clipping/ms/correlation share one pass.

const band = chs => ({ lo: chs.map(lowEnergy), hi: chs.map(highEnergy) })
audio.stat('lo', { block: band, reduce: rMean })
audio.stat('hi', { block: band, reduce: rMean })

A record can also carry fields no stat is named after: list them in extra and they land in a.stats beside the stat's own. A field that holds its block's place on a grid counted from the start of the stats names the grid's period: sr => samples: an edit that moves blocks by other than whole periods drops the field, and a query that needs it measures again (a range: the range itself). The energy stat does this for the 100 ms grid of BS.1770's gating blocks (kcut1, kcut2), so integrated loudness is exact from block stats. a.stats.length is the samples the blocks cover; the last block can be short.

reduce is (blockValues, from, to) → number — it combines the values returned by block, enabling a.stat('mystat') scalar and a.stat('mystat', {bins}) binned queries.

query adds a derived aggregation: query(stats, chs, from, to, sr) → value. Used for stats that derive from other block data (e.g. db derives from min/max, peak from min/max, rms from ms).

a.stat(name, { bins: n }) gives n numbers on the block grid: reduce over each bin's blocks, else query over them, else a stat plugin (or a stat with its own method) over each bin's { at, duration }; a null is NaN. A stat whose value is not a number (events, a key, a spectrum) answers once for the whole range.

ctx has sampleRate and persists across blocks within one decode session (and across the blocks of a playback, for the meter) — set any property for stateful computation.

Registered stats auto-participate in the playback meter — a.meter('mystat', cb) streams per-block values during playback. Block-defined stats emit the raw block value; query-defined stats are evaluated against a single-block pseudo-stats window.

Stat option values that are audio instances (e.g. similarity's ref) pre-render to PCM:

await audio.use('truepeak', 'similarity')
await a.stat('truepeak')                    // −0.4 dBTP (inter-sample, BS.1770)
await a.stat('similarity', { ref: b })

Codecs

Codec plugins — { codec: fmt, test?(bytes), decode?(bytes), encode?(opts) } — extend what audio() can open (header sniffed via test where magic-byte detection draws a blank) and what save()/encode() can write. Every @audio/decode-* / @audio/encode-* package ships its half as an audio.js manifest — halves merge by format name; the bundled umbrellas keep precedence for formats they already serve (streaming decode stays streaming), so codec plugins matter for standalone hosts and formats beyond the bundled set.

Engine-less hosts

Plugins also run without the engine: audio/batch hosts one over a whole signal, audio/stream over live chunks — same param semantics (defaults, automation functions, smoothing), no plan or context.

import { toBatch, toStream } from 'audio/batch'
const compress = toBatch(compressor, { sampleRate: 44100 })
const out = compress(samples, { params: { threshold: -24 } })

Note events

Note-event instruments (voice, poly) take a notes list — the host compiles it to contract §events slots:

await audio.use('poly')
audio(4).poly({ notes: [{ time: 0, midi: 60, duration: 1 }, { time: 0, midi: 64, duration: 1 }] })

Registry

audio.plugins maps name → package; the packages install with audio. Registry ops are instance methods from the start: a.compressor(-30, 8) records the edit and the package loads (dynamic import) at the first render, mapping the positional args onto its params. Registry stats load inside a.stat(name). The CLI resolves names the same way; a sidechain file is key:FILE (ducker key:voice.wav).

audio.import(spec) is how a registry package loads, import(spec) by default. A bundle whose bundler cannot follow a computed import() (a worker has no import map) replaces it with literal imports: audio.import = spec => loaders[spec](). The editor's worker does this with one generated () => import('…') per registry package.

Op plugins:

dynamics compressor · limiter · gate · expander · deesser · ducker · compand · softclip · leveler · transient-shaper · multiband · fet · opto · varimu · vca — denoise dehum · specsub · wiener · omlsa · dereverb · deplosive · dewind · dewow · declip · decrackle · debreath · rnnoise (optional package: @audio/neural-denoise) — effects delay · chorus · flanger · phaser · tremolo · vibrato · autowah · wah · bitcrusher · distortion · exciter · ringmod · freqshift · multitap · pingpong · slew · noiseshaper · lofi · graindelay · stutter · subbass · sbr · rotary · tapestop — reverb freeverb · schroeder · plate · fdn · spring · shimmer — filter moog · korg35 · diode · oberheim · resonator · spectral-tilt · variable · comb · dcblocker · emphasis · deemphasis · derivative · integral — eq geq · tilt · baxandall · dyneq — spatial widener · haas · panner · autopan · midside · microshift · surround — shift pitch-shift · vocoder · formant-shift · paulstretch — color tape · transistor · waveshaper · multisat · amp · cabinet · defeedback — generate osc · noise · chirp · pluck · risset · rhythm · sfx · kick · cymbal · snare · adsr · voice · poly · fm · bell · epiano · modal — more tube · isolate · tune

Stat plugins (land on a.stat(name)):

loudness truepeak · lra · replaygain · dr · speech-contrast · sounds — spectral rolloff · spread · slope · flux · contrast · ltas · zcr — mir structure · tempogram · melody · downbeat · fingerprint · drums · multif0 · transcribe · similarity · coversong · chroma · tonnetz

Beyond the registry — kernels whose inputs aren't scalar params ship as plain packages for direct import: @audio/reverb-convolution (impulse response), @audio/eq-fir (response curve), @audio/eq-crossover (SOS designer), @audio/tune-midi (guide notes), @audio/denoise-repair (regions), @audio/synth-dtmf (digit string), @audio/synth-wavetable (tables), per-band forms of multiband/dyneq/multisat, and the @audio/measure, @audio/sinusoidal, @audio/voice tool/substrate families.