Is this open source? I’ve been working on https://codeberg.org/olpad/openmic which is in the more traditional domain of USB audio interfaces, but I’d be interested in taking a look at the internals of the ETH-68
Yes. And no. The affected users usually have no control over the situation, so if the site owner cares at all that those people can't access it, then it is effectively their problem as they can potentially do something about it.
If neither party cares enough, then it is nobody's problem.
Well, it's certainly a dumb situation that I don't like that's created by my employer (and many others with similar setups), but it takes only two clicks to pick a category to make it available to folks behind these firewalls.
I went ahead and chose "blog" because it sounds like it's a blog. Apparently the "newly seen domains" category goes away on its own after 30 days, so that would have fixed itself, though then it would have been "uncategorized" and still blocked.
As someone who uses a USB equivalent to this piece of hardware (in my case, a Behringer UMC1820): having MIDI and audio inputs on the same interface is a nice convenience feature. Avoids needing to use up multiple ports on the host machine, and avoids a lot of the time synchronization issues that arise with trying to use multiple interfaces at once.
Thank you! You’re right, I didn’t discuss the MIDI implementation in the blog post. Right now it’s a little primitive but functional. Each MIDI message (e.g 3 byte note-on message) maps to 1 UDP packet. I have a simple script on the host side that interfaces these UDP packets with an ALSA loopback MIDI interface.
How does the receive side recover the transmit side's sample clock? There is a BNC for clock sharing between "multiple units", but I'm not sure if that's used / required between transmit and receive.
Or is there no such synchronization, in which case there would be long-term drift?
I think I'm going to have to make a blog post addressing some of the clocking questions that come up. My brief answer for now is that there is only one clock, and that is the ETH-68 clock. The overall data flow is push based; the arrival of a new buffer of data at the Linux host IS the clocking event that schedules an audio graph evaluation. There is no need for another clock on the Linux host.
Unless audio is sampled at ETH-68 clock this won't work for long periods. Except if you are tuning a loop continuously doing continuous resampling using a Farrow filter for example.
That is sort of true, although it should be possible to take advantage of the resampler in PipeWire or JACK (zalsa_in/zalsa_out) to effectively get multiple interfaces in the same graph. I would expect some hit to the audio latency when going through the resampler.
If there were open source Dante that actually worked it would be very cool. Don’t know how much usage it would get since the hardware will always lock it in. But for small studios and research it might be cool. Interesting project nevertheless.
Far be it from me to discourage an H7 build, but I do question the codec choice. It's far from top of the line and it's not like the design is tight on space. TI offers much better (almost 20 dB SNR more on the ADC). Maybe a gen 2 could benefit from a better codec.
Yes this is an older codec but it has been reliably in production for around 20 years, and has a high channel count to cost ratio. The specifications are not cutting edge but I get comparable performance to my trusty Saffire Pro 40.
I have my eye on some other codecs for the next project.
I'm curious what you mean about the STM32H7. My experience has been that it is a challenging part, at least partly due to bugs in the HAL. This is especially true for the ethernet implementation. But once it is working the performance is quite good.
Oh just the I used it in uni, made a design with it (though never assembled it-parts are still in a bag), and I've used it professionally for some years. The HAL certainly has bugs but they were easy to work around for me. That said, I've never used the Ethernet peripheral on it. And everything else has been easy enough for me to go direct register access if needed.
It's a hardware audio interface with a very low latency (buffer of 64 samples, 3.6 ms), which supports both PipeWire and JACK over Ethernet, as shown on the pictures in the post. It's audio over CAT6, and a lot of it, 6 inputs + 8 outputs (TRS balanced) at 48 or 96 kHz.
Imagine having this on the stage right next to your analog gear, and a computer 50m away.
(Opening the page under discussion was actually helpful, it lists all this right at the top.)
Afaik all pro audio standards more or less require ptp for timing, and that might add requirements for nics and switches? otoh the i210 nic, that is mentioned in the post, does have full (g)ptp/phc/tsn/etc support
I currently have a setup with a Raspberry Pi (alpine+pipewire+dac) to stream audio over the network.
This enables me the watch video with no latency problems, because it's embedded in the audio stack (buffered). That's why I was wondering which problem gets solved with this project.
But now I understand that this is for concerts not for some audiophile multi room setup.
Pro audio systems frequently (and increasingly) use networked audio. Some obvious uses are distributing the sound from the instruments and people on a stage over to the front-of-house mix position that's usually somewhere mid-crowd, and also to the monitor mix position that's usually in a vaguely-quieter area off to the side of the stage, and to the broadcast truck.
The old tried-and-true method also still works: Analog splits. Take a bunch of audio sources (eg, microphones) and plug them into passive stage boxes that output over a thick-ass cable. Those thick-ass cables go to larger passive split boxes (often on wheels by this point), with two or more outputs for even-thicker cables, with one pair of wires for every individual signal -- often with individual shielding and jacketing.
Eventually, these splits can deliver audio to the different places that need it -- where it's ultimately broken back out into a bazillion individual cables that get plugged into things like mixers.
The cables can be very long (hundreds of meters) in length, and extremely heavy. They're expensive to produce, they're expensive to maintain, and they're expensive to wrangle. They often get transported in their own dedicated wooden trunks. But at least it's simple: A bunch of different audio devices scattered all over a venue, wired in parallel, listening to the signals that are directly produced by microphones on a stage.
---
But with networked audio, it can be more like this: A few boxes on a stage that accept analog audio on one side and emit network frames (often Ethernet or Ethernet-adjacent, and actually using IP isn't a rule at all) that contain digital audio on the other side. Those frames go to a network switch. One or more tiny-ass network cable comes out of the switch and goes wherever it needs to go, and switches can be cascaded, and more audio channels can be added downstream. Because Ethernet(ish) is a many-to-many network, it's bidirectional, too: Audio signals can go upstream just as easily as they go downstream.
It's tidy. It works. It's still expensive because the endpoints are expensive, but the cables themselves can be fairly inexpensive (think robustly-built Cat6 or fiber patch cords instead of giant cable trunks). If the venue's infrastructure goes to the right places and can be trusted, then it can also be used: Plug the stuff from the stage switch into a fiber patch panel on the building, and plug the broadcast truck outside into the same building, tie them together in some MDF or IDF somewhere, and send it. (And in a pure and just world where dedicated fiber links both exist and are easy: Patch another into the studio downtown. Or route it over an IP link that is shared with other purposes, if appropriate and also feeling brave.)
But with the tidiness comes complexity. Like... Latency is kind of a big deal here in ways that aren't a practical issue with analog audio. Putting too much delay between a vocalist and the monitors that they hear themselves with is actively deleterious of their ability to sing, for example.
And buffers are still required (they're ~always required when packet-switched network frames get converted to continuous analog signals). Keeping the buffers small requires very tight timing signals that get shared between all points. Pre-existing systems often achieve that with things like Precision Timing Protocol (though variations exist).
That all conspires to mean that the heavy lifting at the endpoints is often in the realm of FPGAs.
But, again: We get many channels over some bog-standard network cabling. Dante, for example, can be used to transport hundreds of 48KHz 24-bit audio channels on one gigabit ethernet link.
---
Anyway: This is a cheaper, smaller method. It uses an STM32H7 microcontroller to convert betwixt the network transport stuff and the DACs and ADCs of the analog world. It's designed to be used with the open-source Jack system that is commonly-used internally whenever Linux gets involved in recording or stage use, so it's simple to integrate with a Linux PC running software like Reaper. And at the end of the day, it's transportable over the Ethernet networks we all have.
And despite being built around an STM32, it achieves quite usable latency: The stated 3.620ms is about the same as a 1.2 meters of distance for sound in air.
Neat stuff. I'll probably never use it, but it's neat. :)
That's may be how the back-end of a modern broadcast FM radio transmitter works, but the good old stereo radio transmission itself still the same as it's been for many decades: It is analog, and mid-side encoded, and it remains completely compatible with monophonic receiver implementations.
Stereo pair, pilot, RDS (the scrolling artist/song text), and the 67 kHz SCA. It's not quite 192kHz as I recall, but that's nearest fit really. AES192/192kHz MPX composite on the back end of every exciter/transmitter when I was last near them.
That's for the FM baseband signal, not the audio. Among other things you're doing here, you're basically treating a stereo signal (plus a pilot that effectively contains no information, plus an extremely low bitrate RDS stream in an extremely inefficient way) as one monaural signal.
But maybe there's a use case for replacing an AES192 signal carrying the full FM baseband signal over ETH-67? I'm not sure it's the correct fit, though.
That's how radio works now, it's owned by like 3 companies. Automation/playout for large groups of stations is all centralized, and they carry the baseband over IP to transmitter. Radio solved it awhile ago.
Distributing music != recording music != processing music.
There ARE good reasons for recording and processing music at higher bitrates and sample sizes. Effects, especially those with positive feedback loops, can go more unstable and clip and lose information with fewer bits and samples.
There's just no good reason to DISTRIBUTE music at the higher rates.
Don't such effects already upsample and downsample as needed internally ? You don't need to waste cpu on the effects that don't need higher sampling rate, right ?
They could. But then you get clipping effects and sampling effects when you convert back and forth.
Better to upsample once (generally a non-lossy operation) and operate on single precision FP (24-bit mantissa and an exponent) at higher sampling from that point forward and then downsample once (generally a LOSSY operation).
Sorry, I was reading this late at night and must have misread the microcontroller. Then, the limitation lies in the codec and it would not be a minor modification or a matter of trading bit rate for latency.
I need to record ultrasound for work onboard construction vessels, currently we use either specialized equipment for PAM, which is limited in some aspects, or a USB sound card with a SBC and a hacky setup for sending PCM over TCP.
I would love this and am actively looking for a unit for my linux setup but: why the limit on sample rate? why not also support 44.1 kHz ?
How am I supposed to master for CD, which is still something people do?
Great to hear that you will support it in the future!
Sure, resampling works, but i only want to do it when it can't be avoided, not for basic playback.
If this had open source firmware, I think I know a lot of people who would be interested, but I can't seem to find out if that's true or not and where to buy one.
Thanks for your interest! I made a single reddit post about ETH-68 this week and this has been copied around various forums. My intent was to figure out if anyone thought this would be cool enough to produce. I'm trying to figure out if I should do a production run, open source it, or some combo of the two.
I'm currently running 4x Echo Audiofire 12s in my Linux based Studio and something akin to what you are doing would be a much more fun and less expensive option than RME. Next version is 24 channel?? ;)
I think this would be a cool project to do as a DIY.
The market for a linux native Audio Interface is there I think.
At least I for one would appreciate this!
> The typical default latency for a Dante audio device is 1 msec.
(emphasis mine)
The latency depends on the device. Hardware implementations of Dante commonly support latencies of 1ms or less, but software implementations are higher. The minimum latency of Dante Virtual Soundcard running on a PC is 4ms.
That's the appropriate number to compare against here (since the PC is using a software driver to interface with the network). However, that 4ms number is one-way latency, and the OP's 3.6ms number is round-trip. So this is already half the latency of DVS. (That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.)
> That being said, it sounds like this latency figure was only achieved in very ideal configurations, and we don't know the reliability/rate of late packets compared to DVS.
Hm, the audio latency doesn't fluctuate so it's not like the testing conditions affect the measurement. The audio latency is a fixed quantity that depends completely on the number of storage elements in the data path which isn't variable.
Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting. The histograms on my page (which should be animated BTW) show real time processing latency measurements while running the audio all the way through Bitwig with a moderate DSP load (multiple instances of Pianoteq, samplers, live MIDI input). On my system (details at the bottom of my page), I can do this at 48 kHz and 64 sample buffers with zero underruns. If I drop down to 32 samples per buffer, I do start getting underruns.
All I can do from the hardware side is try to minimize the processing latency of a typical cycle so that there is more head room to absorb jitter. The vast majority of the jitter comes from the Linux host. It's up to the end user to tune the system for low jitter. This is usually the case for audio on Linux, and the rabbit hole can go pretty deep on system tuning.
> Perhaps the better thing to focus on is the frequency of underruns (or "xruns" as they say on Linux, which also covers overruns) for a specific sample rate and buffer size setting.
Right, that's what I meant. There will be some jitter depending on the scheduler, system load, network performance and traffic, and the quality of hardware/drivers; so more latency gives you headroom to absorb the jitter without underruns. The "ideal configurations" I was referring to were the low-traffic network and high-quality NIC. On a setup with more jitter, you might have to increase the buffer size (and thus latency) for reliable operation.
Mostly, I was just trying to contextualize the numbers for readers who aren't super familiar with low-latency audio networking. Sure, this project may not achieve the sub-1-ms roundtrip latencies that you can get with dedicated Dante hardware (like a Yamaha mixer and stagebox); but Dante can't do better than 8ms when one of the ends is a PC (although they were probably aiming for reliability on setups not aggressively tuned for minimum jitter.)
gigabit have throughput, but not latency. For small packets and on low load link there will be almost no difference (i had better source: https://serverfault.com/questions/276651/network-latency-100... , ethernet cnc controllers have same problems). It needs to go 10gbps or even 25gpbs to see lower latency, where electrical signal uses much higher electrical signaling bandwidth.
But the packets aren’t small when streaming audio. The latency improvement is substantial for 1000 byte packets as the accepted answer in your SO post demonstrates
Ah, then my blunder. I saw comments about 64/128 sample buffers automatically tough that most packets will be in 128/256 bytes range where difference not that big.
1G would certainly decrease the processing latency (not the audio latency) by quite a bit. STM32H7 doesn't have a 1G MAC. 1G MAC is kind of rare on "friendly" microcontrollers, although there are at least two that I'm evaluating for the next project.
It seems like, at 16000TbaseT anyways, you’re adding an overhead of about 150% on top of copper/fiber latency plus transmit time latency to process data at that bandwidth. I wonder if the same holds true at 100 vs 1000, 2500, 10000? Certainly this is a known tradeoff for DDR performance tuning — if you don’t mind spiking response times greatly, you can get the advertised maximum speeds, else you accept less bandwidth for somewhat less latency — and they’re both effectively using the same strategies to talk over copper.
Lol, this wasn't my choice really! The JACK server's default UDP listen port is 3000 with the netone backend, so that is default port that ETH-68 sends to. This can be adjusted on both ends
The choice of port wasn't the complaint so much as the choice of IP address. 12.12.12.10 is a real IP address on the Internet, so if you plug this device in somewhere with an Internet connection, it will begin spamming nuisance traffic to some random device out there. It would be better to choose an IP address in one of the private spaces for this purpose.
As the owner of the PCIe card that measurement was taken from, yes, I’d say it’s quite impressive. The average round-trip latency for a USB audio interface at 48 kHz/128 samples would usually fall somewhere around 8 ms, which is a bit much if you’re monitoring post-FX.
On Linux, we don’t have much in the way of Thunderbolt support for audio interfaces, so the only way to achieve this sort of latency has traditionally been with PCIe or PCI audio interfaces. Having a low-cost, infinitely more portable solution would be very welcome.
They quote a roundtrip of 3.6ms, so one-way 1.8ms. In the plots on their page, it looks like the processing latency is centered on around 1ms:
> The LATMON pulse width is therefore an accurate measure of the total processing latency of each cycle and is affected by every element in the data path: processing delay in the microcontroller, network transmission delay, host OS delays, signal processing delay in the DAW, etc
You are confusing the audio latency with the processing latency. I see this mistake a lot.
One way audio latency is about 3.6 ms divided by two = 1.8 ms. This type of latency is audible.
Processing latency is just the amount of time it takes to complete all processing for each audio cycle. From the histograms, the typical processing latency is about 625 us and of course there is some jitter (almost all of the jitter comes from the Linux host BTW). Processing latency is not audible. However if the processing latency exceeds the deadline on a given audio cycle, there will be an underrun which will cause an audible glitch.
reply