r/DSP

▲ 4 r/DSP+3 crossposts

ZedBoard PYNQ image?

I’m working with a ZedBoard and Vivado v2024. I can only find older PYNQ images for the ZedBoard.

What’s the recommended way to build a newer PYNQ image with the required drivers and overlay support?

If anyone’s done this recently or has a working build setup, I’d really appreciate some pointers 🙏

reddit.com
u/No-Statistician7828 — 1 day ago
▲ 150 r/DSP+9 crossposts

asm.fm — a chiptune synthesizer in pure x86-64 assembly (no libc, no audio lib)

Learning project that became my favourite: a chiptune synth written entirely in x86-64 assembly (Linux, NASM). No libc, no audio library — just computing raw 16-bit samples and writing a WAV header by hand.

The premise is that sound is just a list of numbers (44100/second) describing where a speaker sits. So the whole synth is: generate the numbers, write them out.

It does four oscillators (square/saw/triangle + LFSR noise), polyphony by mixing voices into one buffer, ADSR envelopes, and FM synthesis with a hand-built sine table. Working on effects next (vibrato, delay, reverb).

github.com/whispem/asm.fm

Feedback on the low-level details welcome — especially the fixed-point math in the FM operator.

u/whispem — 4 days ago
▲ 17 r/DSP+1 crossposts

Analyzing musical concert seismics

Hi

So we did a fun thing and recorded a concert in the university with geophones and are now analyzing the data (1000-3000 people, outdoors, around 20 triaxial geophones planted in grass in a around the venue (see picture), 500 Hz sample rate, velocity model should be 300-500 m/s).

Our biggest hopes and dreams were separating crowd movements from PA and music but we're still very far from that.

Right now we're trying to detect song starts and BPM using both STALTA and spectral analysis, similar but not exactly like that one Taylor swift concert paper and trying to use song starts to measure correct distance between geophones in hope to locate the major signal sources (stage PA and crowd mass).

  1. STALTA results show major time delta between near geophones (scale of 0.05s) and sometimes geophone closer to the stage see the first peak significantly later than further ones. How would you make sense of these results? (sta: 0.1s, lta: 10s)

2.. An open question for anyone interested - how would you go about this? Which tools and algorithms would you use?

  1. To detect BPM we tried to autocorrelate single geophone traces checking the lags that match musical bpm (50 - 180 bpm so 0.3 - 1.2s) and locating the highest peak in the range - some songs were a dead match so the studio version bpm, some were half or double and some were wrong, we haven't gone through the meticulous process of using all geophones yet but did find that sometimes the entire spectrum yielded a more accurate result and sometimes focused bands under 50 Hz were best.

  2. What is the best way to differentiate PA and crowd-based signals in a rather noisy recording? especially when the crowd was much smaller comparative to previous works in the field (the geophones were also closer in our case).

  3. Do you think paved roads between the signal sources and all geophone and terrain height differences between the stage and some geophones are a significant hindrance?

Geophone geometry: before you come for us about the geometry keep in mind that like every good field work, we had some issues with production and the law and had to compromise our array.

https://preview.redd.it/zcsszw2c26kh1.png?width=1714&format=png&auto=webp&s=1c97b8d1144c9d8deea81960c1499e3283d2214e

Thanks to anyone that read all the way!

reddit.com
u/avengerbob147 — 3 days ago
▲ 9 r/DSP

Which metric for audio quality measurement is better?

Hello, I'm developing an audio codec, and currently stuck on one thing: which metric should I use for quality measurement?
For now I've been using gstPEAQ Basic model (because Advanced really sucks speed and correlation wise). I was thinking about ViSQOL, but their results also kinda suck, especially for lower bitrates like 64 kbps, where difference between 64 kbps OGG and my codec are really noticeable, and yet differentiate only around 0.1 MOS-LQO (gstPEAQ Basic makes difference big, around 2 ODG).
Any suggestions?

reddit.com
u/Alfoser — 2 days ago
▲ 0 r/DSP+2 crossposts

Have you ever thought about what timbre physically is?

When you hear an instrument, you can tell it's a piano, a violin, a drum — even if they're all playing the exact same pitch. We call that difference "timbre." But it made me wonder: can you actually explain what timbre is, physically?

In school you probably learned that pitch comes from frequency — how many times the air vibrates per second. But since the same pitch can have different timbres, frequency clearly isn't the answer.

If you've messed with synthesizers, you might guess it's waveform shape — sawtooth vs. square vs. pulse. That's closer, but still not quite right either: waveforms that look completely different can sound identical.

I dug into this and wrote it up (free, no signup) with interactive audio demos so you can test it on your own ears.
Link: https://sciencemusic.pages.dev/timbre/

This post was created using AI assistance (mainly translation).

u/LowFaithlessness426 — 4 days ago
▲ 39 r/DSP

You can date a TR-808 from a recording by fitting the shape of its TONE knob

I am modelling a TR-808 from the service notes, stage by stage, and while doing the snare I ran into something I did not expect: two independent ways to tell which production revision a recording came from, one of which is just the response curve of a potentiometer.

Background. The 808 snare is two bridged-T resonators in parallel, one for he shell body and one an octave up, mixed by the TONE control, with the noise ("snappy") path summed in separately. Mid-production Roland changed four parts in that voice. The service notes list them as a PORTION CHANGED

annotation:

C58 .027 to .056

C61 .0068 to .015

R191 8.2k to 10k

R200 47k to zero

The caps are the obvious half. Only one cap of each bridged-T pair changes, so the frequency ratio is sqrt(27/56), not 27/56:

resonator 1: 249.63 Hz -> 173.34 Hz (-6.29 semitones)

resonator 2: 499.00 Hz -> 335.98 Hz (-6.85 semitones)

So an early 808 snare sits about a fifth higher, and because the two shifts differ slightly, the pair goes from an exact octave to a flat one. Q drops a little too (17.4 -> 16.3 and 10.7 -> 9.9), so the shell also rings around 35% longer after the change.

Now the part I thought was neat. R200 sits in series between an op-amp output and one end of the TONE pot. Since 47k is comparable to the pot's 100k, it is not a level trim: it changes the SHAPE of the crossfade law across the whole sweep. Deriving both cases:

R200 = 47k -> TONE spans 16.8 dB

R200 = 0 -> TONE spans 26.8 dB

The recordings I have span 26.4 dB measured. Fitting the derived law against ten measured amplitudes with a single common gain and nothing else fitted gives 1.16 dB mean error for R200 = 0 against 6.62 dB for 47k. That is a clean verdict, and it identifies the machine's revision without looking at the resonator frequencies at all. Then the frequencies agree independently: measured 172.0 and 339.5 Hz against derived 173.3 and 336.0 for the late version.

Two consequences that might be useful to others:

  1. A control law is a fingerprint. We normally treat pots as "the user's problem" and model them as a gain or a simple taper. But when a series resistance is comparable to the pot's own resistance, the law is a property of the circuit, it is measurable from audio, and it carries information the spectrum does not. I would not have thought to look if the two cap values had not sent me to check what else changed with them.

  2. Watch out for manuals describing a different machine than the one in front of you. Roland's own printed spec quotes 476/238 Hz for this voice, which is the early revision. Earlier in the same project I found their printed bass drum tuning (56 Hz, and 62.5 Hz in the prose) disagreeing with both my derivation from the component values and the hardware, which both land at 49.5. Component values have been right every time so far;

printed prose has been wrong twice.

Method note, since it is the part that actually matters: I derive each stage symbolically from the drawn component values, then measure real hardware recordings, and I only trust a number when the two agree. Where a published academic derivation of the same circuit exists I use it as a third reader, and that has caught mistakes in both directions, including three errata in the paper.

Disclosure: I make a drum machine plugin and this work goes into it. Not linking it, this is about the circuit.

reddit.com
u/BeatForge_Dev — 4 days ago
▲ 21 r/DSP

What are some interesting and niche fields where SP is applied ?

Aside from communication, imaging and audio processing, are there other STEM fields that have been using SP tools to solve their problems?

reddit.com
u/al3arabcoreleone — 5 days ago
▲ 3 r/DSP

Physics undergrad seeking advice

I'm a rising sophomore at a top liberal arts college who's planning on going into some form of music technology in grad school. I have been on a physics major track and I'm looking at DSP as one of the possible areas to explore -- any suggestions on what classes/opportunities I should consider for me to be better prepared in the field? Or, if there is any other niche you think might be worth looking into, I'd greatly appreciate the response. Thank you!

reddit.com
u/Pristine_Ant_6817 — 5 days ago
▲ 16 r/DSP+6 crossposts

Rewrote the interpolation on my chorus after getting corrected

I posted my chorus plugin here a couple weeks ago and someone told me allpass interpolation was the wrong tool for a modulated delay, since the recursive state goes stale while the delay length is moving under it, and that I should look at Lagrange. I was still on linear at that point. So I wrote a 4-tap cubic Lagrange interpolator by hand instead of pulling in a crate, mostly because I wanted to understand it rather than trust it.

What's actually in the thing now, so nobody has to guess from the post:

- 6 ms base delay with the LFO modulating around it, cubic Lagrange on the fractional read. Taps at -1, 0, 1, 2 off the integer index, each one wrapping the ring buffer separately.
- One-pole highpass at 160 Hz sitting inside the feedback loop so the low end doesn't stack up.
- One-pole lowpass on the wet path, cutoff swept by the LFO between 1k and 7k. Most of the character comes from that, not the delay.
- Separate voice struct per channel, LFOs offset 90 degrees. Collapses to mono without eating itself.
- Dry/wet is a plain linear crossfade. There's roughly a 3 dB dip at 50% and I'm fairly sure that's comb cancellation from summing a correlated wet signal with the dry, not bad crossfade math. If that reasoning is wrong I'd like to know.

One debugging note that saved me: check your coefficients at frac = 0.5. They should land on -0.0625, 0.5625, 0.5625, -0.0625. Mine didn't, and it was a typo in one of the products in the third coefficient.

At my defaults (1.6 Hz, about 4 ms depth) the improvement is small. A little less grit on held notes up top. I assume it opens up more if you push the rate.

Two things I still can't answer:

  1. Any reason to go past cubic when the modulation is this slow, or is higher order mainly for pitch shifting and wide sweeps?
  2. The HF droop moves with the fractional part, so slow modulation means a slow wobble on the top end. Does anyone correct for that in practice or is it under the floor at this depth?

It's free, GPLv3, VST3 and CLAP, mac/windows/linux. It's built for pitch-corrected vocals specifically, which is the only thing I use it on. If you'd rather hear it than read about it, it's at pyfessional.tech, the download button picks your OS for you and the source is linked from the same page. Would rather have someone install it and tell me it sounds wrong than get upvotes.

u/kiwiberrydrink — 8 days ago
▲ 9 r/DSP+1 crossposts

I wrote a small WebRTC SFU in Go (Pion) because I couldn't read LiveKit

I have a small chat app with voice rooms. Started on LiveKit, but every time the audio broke I had no idea what was going on inside. Too much code for me to read. So I wrote my own SFU on top of Pion. It's been running my rooms since July, up to 25 people.

Just pulled it out into its own repo: https://github.com/Amesu-afk/tarnmedia

About 1800 lines with tests. No simulcast, no transcoding, no recording. It forwards audio and video, checks JWTs, and that's it. Every subscriber gets the publisher's single encoding, so one bad connection makes the room worse for everyone. I know. That's the next thing I want to fix.

That's also where I'm stuck. If you've added simulcast to something this small, how did you keep the packet path from turning into a mess? Right now the forwarding code is simple enough to read in one sitting and I'd like to keep it that way, but layer selection looks like it touches everything.

u/am3su — 8 days ago
▲ 263 r/DSP+2 crossposts

Totally Accurate Audio Frequency Spectrum Chart

u/novateai — 9 days ago
▲ 2 r/DSP

How to simulate spectral leakage by hand on paper?

say I want to simulate by hand with a dummy signal and show how spectral leakage will occur. How shall I do it?

Spectral leakage is a spread that is present across the entire frequency spectrum caused by the rich harmonics generated by the aperiodic samples.

And windowing eliminates spectral leakage. Windowing means to multiply N sampled signals by a window of the same length to remove the discontinuities at the edges. There are various types of windows like rectangular, hamming, hann etc.

Now my objective is to take a signal, and with pen-and-paper simulate the rise of spectral leakage, and then how windowing eliminates it. Help me achieve it. I read many books but they do not provide such analysis.

reddit.com
u/Embarrassed_Grab6901 — 6 days ago
▲ 18 r/DSP+1 crossposts

How to send a 500MHz signal via SMA using GTX Transceiver on Zynq-7000 (XC7Z035)?

I am working with an AMD Xilinx Zynq-7000 SoC (XC7Z035) board and need to send a 500MHz signal through the board's SMA output. The problem is that we are using the GTX Transceiver Wizard to generate this signal, and it is proving to be extremely complex.

I could barely find any tutorials on YouTube about it, the Xilinx manual is also quite hard to follow on this point, and there are almost no articles or forums discussing this specific use case. This has left me wondering whether the GTX Transceiver is simply not the most common path for this kind of application, since the scarcity of material suggests that few people use this flow.

I am hoping to hear from anyone who has been through this. Is the GTX Transceiver really necessary to generate a 500MHz signal on this board, or is there a more direct path, for example through some clocking IP or an external DAC? Does anyone have a reference, tutorial, or example project showing this GTX Wizard configuration flow in practice? Is there an alternative connector or approach you would recommend for sending this signal to the oscilloscope without relying on the GTX?

Any pointer, link, or even a suggestion that I am overcomplicating this would help a lot. Thanks in advance.

u/Matheeodua — 8 days ago
▲ 27 r/DSP+3 crossposts

After a year of JUCE/C++: BeatForge — REX player + 808/909/303 engines, all DSP hand-written

Been heads-down on this for about a year and figured this crowd would appreciate the guts more than the marketing page.

BeatForge — drum machine, step sequencer and REX loop player. VST3/AU/AAX/Standalone, universal binary. JUCE 9 (migrated last month), everything below the framework is my own code.

Bits that might be interesting to people here:

REX SDK. Wrote a JUCE wrapper around Propellerhead REX SDK 1.9, including drag-and-drop of slices out to the host. Bitwig still has no REX support, which turned out to be a decent reason for the plugin to exist.

Synth engines. ~5k lines of 808/909/modelling done from schematics rather than sample playback — VCA topology, click transients, pitch sweeps, the 909 tom triple-VCO, that sort of rabbit hole.

Antialiasing. Started with blanket 8x oversampling and it was eating CPU for no good reason. Replaced it with ADAA on the memoryless waveshapers (exact tanh with the ln(cosh) antiderivative) and linear-phase minBLEP saws at 1x, then a per-engine oversampling map — most engines run 1x now, only FMKick and the 909 snare need 4x. Built an offline FFT alias-measurement harness to verify it; alias floor sits around -90 dB. Worth noting ADAA-2 blew up on me under heavy drive — catastrophic cancellation, auto-mute and crackle — so everything is bounded ADAA-1.

Realtime discipline. Lock-free atomic slot for deferred time-stretch, ScopedAudioSuspend guards on live mutators, steal-quietest voice allocation. Debug assertions went from ~3150 to 3. auval and pluginval at strictness 10 both clean.

Also. Full undo/redo, NAM master-bus saturation, and a WASM build of the synth engines for the site via Emscripten and a hand-rolled juce_shim.h.

Happy to go into detail on any of it — the ADAA and REX wrapper work especially, since there's not much written about either.

Site + demo: https://www.beatforge.nl

u/BeatForge_Dev — 12 days ago
▲ 15 r/DSP

DAWG - Digital Audio Workstation Game - Free Public Beta

Hello DSP crowd!

After around 10 months of building DAWG - Digital Audio Workstation Game, and roughly four months since the last public test, I have finally uploaded a new beta.

It is now much closer to my original goal: a real music making system with a game built around it, rather than a game with some DSP bolted on.

DAWG tuning panel & visualization

The audio engine is custom DSP written in C# and compiled with Unity Burst. It includes:

  • Subtractive, FM, and wavetable synthesis
  • Per-instrument DSP chains
  • Send/return and live performance FX
  • Real-time parameter control and MIDI input
  • Custom live visualization following the signal through the complete DSP chain
  • Cross-platform multiplayer synchronization of the clock, patterns, and DSP parameters

Since DSP is the foundation of DAWG, I would really appreciate feedback from professional public.

I am especially interested in opinions about the sound quality, oscillator and filter behaviour, aliasing, FX routing, parameter ranges, stability, and anything else that feels questionable or incorrectly implemented.

The current public beta is free to download on Itch and you can also leave a rating directly there.

Download: https://dawg-tools.itch.io/dawg-digital-audio-workstation-game

Please do not hold back just because it is an indie project. I know there are still rough edges and technical criticism is exactly what I am looking for, but I think DAWG is ready to meet professional audience.

Thanks for your time!

reddit.com
u/Emotional-Kale7272 — 12 days ago
▲ 7 r/DSP+2 crossposts

I built a local macOS/Windows tool for detecting AI-like artifacts in finished music

I have been working on AI Track Inspector 2.3, an experimental offline application that estimates whether a finished track contains acoustic patterns commonly associated with AI-generated music.

This is not intended to prove authorship or identify a generator with certainty. The output is a calibrated classifier estimate accompanied by temporal coverage, model agreement and reliability information.

How it works

The audio is converted to a normalized 16 kHz mono analysis signal and divided into overlapping four-second windows with a one-second hop.

Each window is evaluated by several branches:

  • A spectral “fakeprint” branch looking at frequency-domain texture, high-frequency structure, spectral flatness, roll-off, transient behaviour and related statistics.
  • A log-mel convolutional neural network trained on time-frequency representations.
  • A rhythm branch measuring beat-grid consistency and local timing behaviour.
  • A fusion stage combining the branches while accounting for disagreement and out-of-distribution input.

The application then produces:

  • An AI-likeness classifier estimate.
  • A separate timeline showing where AI-like acoustic patterns were detected.
  • Temporal coverage across the track.
  • Reliability and branch-agreement information.
  • An ambiguous verdict when the analytical branches strongly disagree.
  • PDF and JSON reports.

The training and evaluation material included Suno and Udio tracks, mastered and pitch-shifted AI examples, commercially produced real music, demos, semi-live recordings and difficult human-made negative examples.

One important change was removing absolute or unusual BPM as independent AI evidence. Speed-ups, half-time/double-time interpretations and unreliable tempo tracking produced too many misleading results. BPM is now hidden when the beat grid is uncertain and rhythm can only contribute when supported by the other branches.

Implementation

The macOS version is built with SwiftUI, AVFoundation, Accelerate/vDSP and Core ML. The Windows version uses a local Edge/Chromium audio runtime with the same model weights and fusion logic. All processing happens on the user’s computer; audio is not uploaded to a server.

The detector is designed for complete tracks and final mixes, not isolated stems. Stem-level results can be misleading because their spectral and temporal distributions differ substantially from full arrangements.

I am particularly interested in feedback about:

  • Codec and resampling robustness.
  • Cross-platform decoder differences.
  • Better out-of-distribution detection.
  • Hard-negative dataset design.
  • How uncertainty should be presented without turning a model score into a false claim of certainty.

The current builds are available here: https://www.dropbox.com/scl/fo/bo5dd9t04f818jc8uos44/AOp5_gUP0L6uc63o0nNXcRs?rlkey=9i817ml9fh7mkk0y6ebxaoq8v&dl=0

It is a free experimental project, and I would appreciate technical criticism, difficult test cases and ideas for improving the validation methodology.

u/rootsashok — 11 days ago