Competing Modulators — User Guide
A deterministic multi-voice FM generator in which three sinusoidal control rates compete inside each voice's bounded instantaneous-frequency trajectory.
What this does
Competing Modulators generates sound from scratch; it does not process an input Sound. It creates between two and eight sinusoidal carrier voices. Each voice has its own carrier frequency and its own three-rate modulation pattern. The three modulators are combined into one bounded control signal, that control signal moves the voice's instantaneous frequency, and the resulting frequency is integrated at the audio sample rate to obtain oscillator phase.
The voices are then mixed in mono or positioned independently in stereo, followed by one optional global envelope, a short edge fade, and optional peak normalization.
What “competing modulators” means
The word competing describes the control structure, not a physical interaction between separate synthesizers. For each voice, three sinusoidal modulators pull the instantaneous frequency in different directions at the same time:
The weights sum to 2.15, so the combined control m(t) remains bounded between −1 and +1. The three rates are related but deliberately not identical:
Because these rates continually move in and out of alignment, the summed control develops beating, reinforcement and cancellation. Different voices use different carrier frequencies and different rate sets, so their FM sidebands also overlap and interfere in the final mix.
Quick start
- Run the script with no input Sound required.
- Choose a preset. Gentle Interference is a good starting point.
- Choose the number of voices and the modulation intensity if using Custom.
- Use Modulator base rate to set the basic speed of the control motion and Modulator spread to separate the competing rates.
- Choose a global envelope and a spatial mode.
- Run the script. The resulting object is named
competing_<preset>.
Signal model
1. Carrier layout
Voice 1 starts at Base frequency. Each following voice is 10% higher than the previous base step:
With four voices, for example, the carriers are 1.00×, 1.10×, 1.20× and 1.30× the entered base frequency.
2. Bounded frequency modulation
For each voice, the normalized three-modulator sum controls instantaneous frequency:
d is Modulation intensity. Because m is bounded to ±1, an intensity of 0.50 means a maximum instantaneous-frequency deviation of approximately ±50% around that voice's carrier. At the allowed maximum of 1.0, the lower bound can reach 0 Hz but does not become negative.
3. Audio-rate phase integration
The script does not substitute a time-varying frequency directly into sin(2π f(t)t). Instead it integrates the instantaneous frequency sample by sample:
This distinction matters: it produces a genuine frequency trajectory whose derivative is the requested instantaneous frequency.
4. Voice level
Later voices are progressively quieter, using a 1/√voice weighting. The complete set is then normalized by the square root of the summed voice-weight energy and multiplied by 0.65. This keeps the overall expected voice energy comparatively stable as Number of voices changes instead of simply dividing every voice by the voice count.
5. Practical aliasing guard
FM has theoretically infinite sidebands, so the script cannot guarantee a perfectly band-limited spectrum. Instead it uses a practical headroom estimate: it limits the fastest modulator relative to the sample rate, reserves approximately four times that rate above the largest expected carrier excursion, and automatically reduces Modulator base rate or Base frequency if necessary.
Parameters
| Parameter | Default | What it controls |
|---|---|---|
| Duration_s | 8.0 s | Output duration. Valid range: greater than 0 and at most 120 s. |
| Sample_rate_Hz | 44100 | Output sample rate. Valid range: 8000–192000 Hz. |
| Base_frequency_Hz | 120 Hz | Carrier frequency of voice 1; later voices rise in 10% steps. May be reduced automatically for headroom. |
| Number_of_voices | 4 | Number of independent FM voices, from 2 to 8. |
| Modulation_intensity | 0.5 | Maximum fractional instantaneous-frequency deviation, from 0 to 1. |
| Modulator_spread | 1.5 | Multiplies the second modulator rate of every voice. Valid range: greater than 0 and at most 8. |
| Modulator_base_rate_Hz | 2.0 Hz | Starting rate from which each voice's three control rates are derived. May be reduced automatically for sampling headroom. |
| Envelope_type | No Envelope | One global amplitude shape applied after all voices are mixed. |
| Spatial_mode | Mono | Mono sum or one of four voice-level equal-power stereo layouts. |
| Edge_fade_s | 0.02 s | Independent linear safety fade at both output edges, capped at 20% of total duration. |
| Normalize_output | On | When enabled, scales every non-silent result to a target peak of 0.90. |
| Draw_visualization | On | Draws the mechanism/model/measurement figure. |
| Play_result | On | Plays the final Sound after synthesis. |
Presets
Presets set the main synthesis character by overriding carrier, modulation, voice-count, envelope and spatial parameters. Except for Deep Interference, they do not change the entered duration. They also leave Sample rate, Edge fade, Normalize output, Draw visualization and Play result unchanged.
| Preset | Base | Depth | Voices | Spread | Base rate | Envelope | Spatial |
|---|---|---|---|---|---|---|---|
| Gentle Interference | 100 Hz | 0.30 | 3 | 1.20 | 1.5 Hz | Swell | Mono |
| Metallic Clash | 180 Hz | 0.80 | 5 | 2.00 | 5.0 Hz | Percussive | Wide Field |
| Organic Swarm | 80 Hz | 0.40 | 6 | 1.10 | 0.5 Hz | Tremolo | Rotating Field |
| Digital Warble | 200 Hz | 0.70 | 4 | 1.80 | 8.0 Hz | None | Ping Pong |
| Harmonic Battle | 150 Hz | 0.60 | 4 | 2.00 | 3.0 Hz | ADSR | Stereo Voices |
| Alien Chorus | 140 Hz | 0.90 | 5 | 1.618 | 4.0 Hz | Tremolo | Rotating Field |
| Glitchy Modulation | 220 Hz | 1.00 | 3 | 3.00 | 12.0 Hz | None | Ping Pong |
| Rhythmic Conflict | 110 Hz | 0.50 | 4 | 1.50 | 6.0 Hz | None | Ping Pong |
| Spectral War | 160 Hz | 0.80 | 6 | 2.50 | 7.0 Hz | Slow Fade | Wide Field |
| Liquid Modulation | 70 Hz | 0.40 | 3 | 1.30 | 0.3 Hz | Swell | Rotating Field |
| Crystal Resonance | 440 Hz | 0.50 | 5 | 1.50 | 2.0 Hz | Percussive | Stereo Voices |
| Deep Interference | 55 Hz | 0.60 | 4 | 1.20 | 0.2 Hz | Slow Fade | Rotating Field |
Envelopes
The selected envelope is applied once to the complete mono or stereo mix. It does not alter individual modulators or voices.
| Envelope | Current v0.4 behavior |
|---|---|
| No Envelope | No additional musical amplitude shaping. |
| Percussive | exp(-3t): immediate onset followed by exponential decay. |
| Slow Fade | exp(-0.2t): slow exponential decay. |
| Reverse | Current implementation is a linear crescendo, t / duration. It does not reverse the waveform in time. |
| Tremolo | Amplitude multiplier 0.6 + 0.4 sin(2πrt), where r = 5 + 10 × Modulation intensity; therefore 5–15 Hz. |
| Swell | Linear fade-in over the first 30% of the sound, then sustain at full envelope value. |
| ADSR | One piecewise pass: short attack, short decay, sustain at 0.60, then release near the end. Stage lengths adapt to short durations. |
Spatial modes
Stereo processing occurs at the voice level. The current script does not create width by filtering the completed mono mix.
| Mode | Current behavior |
|---|---|
| Mono | All voices are summed into one channel. |
| Stereo Voices | Voices are distributed evenly from hard left to hard right using equal-power gains. |
| Rotating Field | Every voice follows an equal-power pan trajectory at 0.15 Hz. Voices start at staggered phases around the pan cycle. |
| Wide Field | Fixed equal-power positions from pan 0.05 to 0.95: near the edges, but not completely one-channel-only. |
| Ping Pong | Smooth equal-power movement at 2.5 Hz. Adjacent voices begin 180° apart, so the field alternates rapidly without hard switching. |
Output, reproducibility & level
- Input: none; this is a generator.
- Duration: exactly the requested duration, except when a preset explicitly overrides it.
- Sample rate: the selected
Sample_rate_Hz. - Channels: Mono produces one channel; all other spatial modes produce stereo.
- Name:
competing_<preset name>with spaces replaced by underscores. - Randomness: none in the audio model. A random integer is used only to make temporary object names unique.
- Normalization: when enabled, every non-silent result is scaled to a target absolute peak of 0.90. This is target peak normalization, not an attenuate-only ceiling.
- Raw output: if normalization is disabled and the final peak exceeds 0.99, the Info window prints a warning.
Visualization
The figure separates the mathematical model from measurements of the rendered Sound.
Panel A — Competing Modulators
Voice 1's three exact sinusoidal controls are shown together with their normalized weighted sum. These are model/control curves, not measurements from the output Sound.
Panel B — Instantaneous Frequency
Shows the analytical instantaneous-frequency trajectory for every voice. Horizontal references mark the carrier centers.
Panel C — Model → Measurement
A measured spectrogram of the rendered output with the model instantaneous-frequency guides drawn over it. The guides are control trajectories, not claims that all spectral energy lies exactly on those lines; FM produces sidebands around them.
Panel D — Measured Output
Measured waveform of the final Sound. For stereo results, the channel with the higher whole-file RMS is used as the representative display channel to avoid misleading fold-down cancellation.
The summary strip reports carrier and modulator ranges, FM depth, envelope, spatial mode, final peak/RMS and the practical occupied-top estimate used for headroom checks.
Further Reading
- Chowning, J. M. (1973). “The Synthesis of Complex Audio Spectra by Means of Frequency Modulation.” Journal of the Audio Engineering Society, 21(7), 526–534.
- Roads, C. (1996). The Computer Music Tutorial. MIT Press. See the synthesis sections for broader context on modulation and digital sound synthesis.