That’s very interesting, thank you for sharing. Looks like it could be a very useful tool for testing high performance networking.
I wonder if something similar could be done using TC BPF instead of AF_XDP? My only reservation about AF XDP is that it requires a special NIC to support it, so it may not be useful for a “regular Joe” user. I wonder if TC BPF would also work since it similarly bypasses the Kernel networking stack, I believe you can put packets directly into the NIC TX queue for transmission
tptacek 1 days ago [-]
You can, but the interesting thing about AF_XDP is that you've got a userland path to writing directly to the card's DMA buffers; TC BPF still allocates an skbuff for every packet you send.
adrian_b 4 hours ago [-]
Also with liburing you have zero-copy send and receive operations for normal protocols like UDP or TCP, but this requires a NIC that supports scatter/gather DMA (so that the packet headers go to/from kernel buffers, while the data goes to/from userland buffers).
AF_XDP is also available in older kernels, but with recent enough kernels (zero-copy receive is a recent addition) and with a good NIC, liburing should provide a similar performance.
Palomides 1 days ago [-]
it seems like every NIC on the market that can do 100Gb has support in its linux kernel driver, so probably not a big deal in practice
bgpdude 1 days ago [-]
you can use generic af_xdp which sits at the TC layer. Just get a bit less performance.
ooh no more DPDK, this will make it a lot easier. I just started working with trex but it's all really complicated. Going to give this a try
tptacek 1 days ago [-]
I think, and I'm saying this in part to get someone to correct me, that post-XDP (so 5 years or so now) DPDK is basically obsolete. Is there a circumstance where it would make sense to start from DPDK rather than XDP?
touisteur 11 hours ago [-]
I think access to offload engines is a big part of the appeal of dpdk still, especially for me all the GPUdirect nvidia-only packet steerer.
I need to check about the af_xdp ecosystem around fragmentation/reassembly in UDP too, every time I needed something there DPDK had it, often with an offload path.
Some silly stuff in DPDK are very useful for testing too (in-memory devices).
Also I'm not clear on the virtualization story on af_xdp, with dpdk I got something working at full blast 400G in VMs with little (but finnicky) work.
At this point there's a much bigger ecosystem for DPDK than AF_XDP I think. Also more people know about DPDK than AF_XDP right now e.g. the Ostinato traffic generator's line-rate Turbo functionality uses AF_XDP but most customers assume it uses DPDK.
Disclosure: Ostinato creator here.
trevex 16 hours ago [-]
While I am a big proponent of XDP there are use-cases better suited to DPDK: Being able to offload flows and crypto operations to the NIC is important to a lot of use-cases. The first packet to user space is essentially the slow path even with DMA, that sets up the fast path.
bgpdude 21 hours ago [-]
I think you're correct. Not really aware of any real limitations, other than it's slightly slower than dpdk (it's not a complete bypass), but at a much easier ease of use.
tptacek 21 hours ago [-]
AF_XDP kind of is a complete bypass, right? RX get scooped right off the DMA buffer for the card, and TX get shoved right back in.
bgpdude 21 hours ago [-]
Yeah fair point. I should have been more precise. AF_XDP with ZC bypasses the kernel networking stack and the NIC DMA's directly into UMEM and so in that sense it absolutely is a kernel bypass.
The difference I was getting at vs DPDK is that the kernel NIC driver/NAPI/XDP path is still involvd. With DPDK the userspace PMD is effectively driving the NIC and accessing the queues directly.
either way, it's great and everyone should use it :) that is assuming they have a use case for it. The use-cases are perhaps somewhat limited as it also bypasses the kernel tcp-ip stack, so you gotta do a lot yourself.
I wonder if something similar could be done using TC BPF instead of AF_XDP? My only reservation about AF XDP is that it requires a special NIC to support it, so it may not be useful for a “regular Joe” user. I wonder if TC BPF would also work since it similarly bypasses the Kernel networking stack, I believe you can put packets directly into the NIC TX queue for transmission
AF_XDP is also available in older kernels, but with recent enough kernels (zero-copy receive is a recent addition) and with a good NIC, liburing should provide a similar performance.
I need to check about the af_xdp ecosystem around fragmentation/reassembly in UDP too, every time I needed something there DPDK had it, often with an offload path.
Some silly stuff in DPDK are very useful for testing too (in-memory devices).
Also I'm not clear on the virtualization story on af_xdp, with dpdk I got something working at full blast 400G in VMs with little (but finnicky) work.
A lot of the socket featureset of io_uring seems available in AF_XDP https://docs.kernel.org/networking/af_xdp.html which shows lots of progress since I looked last.
To get an idea of what DPDK gives low-level access to there is the overview https://doc.dpdk.org/guides/nics/features.html and my "favorite annual terabit read" https://doc.dpdk.org/guides/nics/mlx5.html#mlx5-net-features for NVIDIA NICs. Broadcom has some fun stuff too. The first time you hit top RX speed (2x400G my latest) with only one busy core (yay DMA engines) is always a thrill.
Disclosure: Ostinato creator here.
The difference I was getting at vs DPDK is that the kernel NIC driver/NAPI/XDP path is still involvd. With DPDK the userspace PMD is effectively driving the NIC and accessing the queues directly.
either way, it's great and everyone should use it :) that is assuming they have a use case for it. The use-cases are perhaps somewhat limited as it also bypasses the kernel tcp-ip stack, so you gotta do a lot yourself.