Skip to content

#152: Add pure-ibverbs MPRQ raw-Ethernet backend (engine: ibverbs) - #148

Merged
awthomp merged 1 commit into
mainfrom
ibverbs-mprq-backend
Jun 12, 2026
Merged

awthomp merged 1 commit into
mainfrom
ibverbs-mprq-backend

Conversation

@cliffburdick

@cliffburdick cliffburdick commented Jun 10, 2026 •

Copy link
Copy Markdown
Collaborator

Summary

A new ibverbs raw-Ethernet backend (src/managers/ibverbs/), selected via stream_type: "raw" + engine: "ibverbs" and built only when ibverbs is in DAQIRI_MGR. It replaces DPDK's per-packet rte_mbuf alloc/free with a Mellanox/mlx5 Multi-Packet (striding) Receive Queue (MPRQ) driven through DevX, plus a direct mlx5 send-queue TX path. It reuses DPDK rte_ring/rte_mempool for the worker↔app handoff (so it still links DPDK, like the rdma backend) but does not use DPDK for the NIC.

DPDK stays the default for stream_type: "raw"; ibverbs is opt-in for A/B comparison.

What's implemented

Core datapath

  • RX (MPRQ): DevX CQ + striding RQ + TIR with hand-rolled WQE/doorbell management and worker-driven cyclic stride refill; packets DMA strided into one pre-posted MR (host or GPU via ibv_reg_dmabuf_mr).
  • TX: build mlx5 send WQEs directly on a raw-packet QP's SQ (mlx5dv_init_obj, bypassing ibv_post_send) from a slot slab tracked by cyclic index counters; NIC IP/UDP checksum offload + tx_eth_src.
  • GPUDirect, GPU reorder (reuses the src/kernels.cu C ABI), drop/allow-all.

Feature parity work (each its own commit)

  • Multi-queue 5-tuple flow steering (mlx5dv_dr, one domain/table per port) + per-packet flow_id (MARK) via a tag action so flows sharing a queue are distinguishable.
  • Multi-queue-per-core poller — queues sharing a cpu_core are serviced round-robin by one thread.
  • TX checksum + tx_eth_src offloads, get_mac_addr (sysfs).
  • RX hardware timestamps (get_packet_rx_timestamp, CQE TS → mlx5dv_ts_to_ns).
  • print_stats with RX/TX counts + ring-full drop counters.
  • Accurate TX send scheduling (set_packet_tx_time) via a wait-on-time WAIT WQE (HCA-cap-gated).
  • Physical HDS + multi-region on RX and TX — a queue with >1 MR uses a non-striding DevX regular RQ (header→CPU MR, payload→GPU MR); MPRQ stays the fast default for single-region queues.

Validation (ConnectX, cabled loopback, GPU memory)

All hardware-validated; 0.000% PHY loss unless noted.

  • Closed loop vs DPDK (pinned, 64–1024 B): ibverbs 4.75 Mpps vs DPDK 3.9 Mpps; 8000 B 97.5 vs 93.5 Gb/s.
  • Per-core scaling: one poller core sustains ~9.3 Mpps across two active queues (TX or RX); a single core absorbed up to 18.4 Mpps across 4 RX queues on IGX Orin.
  • Per-packet flow_id: two flows (ids 7, 9) to one queue distinguished per packet.
  • RX HW timestamps: consecutive 576 B frames 46 ns apart (exact 100 GbE serialization).
  • Send scheduling: a +2 s hold yields 0 egress in a 1 s run, releases in a 5 s run.
  • Physical HDS: header lands in the CPU MR (dmac/ethertype host-readable), payload in GPU; HDS RX is RQ-depth-bound (~2 Mpps/queue at 128 B vs MPRQ's 17 — the inherent cost of per-packet WQEs, a deliberate opt-in).

Notes

  • All commits are DCO Signed-off-by. Per CONTRIBUTING, the title should be prefixed with the issue number — please advise the issue and I'll amend.
  • Docs are synced: README.md, AGENTS.md, docs/concepts.md, docs/getting-started.md, and docs/api-reference/configuration.md reflect the ibverbs engine and its capabilities (multi-queue steering, physical HDS, etc.).
  • No CI yet (per CONTRIBUTING); validated manually on hardware in the project container.

🤖 Generated with Claude Code

@cliffburdick
cliffburdick marked this pull request as ready for review June 10, 2026 22:49
@greptile-apps

greptile-apps Bot commented Jun 10, 2026 •

Copy link
Copy Markdown
Contributor

Greptile Summary

This PR adds IbverbsEngine, a new pure-libibverbs/DevX raw-Ethernet backend (src/engines/ibverbs/) selected by stream_type: "raw" + engine: "ibverbs". It implements MPRQ via hand-rolled DevX CQ/RQ/TIR/doorbell management, direct mlx5 SQ WQEs for TX, mlx5dv_dr 5-tuple flow steering with per-packet MARK tags, physical HDS via a non-striding DevX RQ, GPUDirect via ibv_reg_dmabuf_mr, RX HW timestamps, accurate TX send scheduling (WAIT WQE), GPU reorder (reusing src/kernels.cu), and multi-queue-per-core polling.

  • The engine vtable, BurstParams free contract (stride-release separated from metadata-pool return), CMake socket→rdma ordering invariant, and compile-definition emission are all correctly maintained.
  • DAQIRI_ENGINE=ibverbs now builds two internal engines (rdma for RoCE, ibverbs for MPRQ); valid user-facing values remain {dpdk, ibverbs}.
  • docs/tutorials/configuration-walkthrough.md line 20 still describes ibverbs as enabling only roce:// endpoints and needs a one-line update for the new MPRQ raw path.

Confidence Score: 5/5

Safe to merge; the engine abstraction, BurstParams lifecycle, and CMake build rules are all correctly maintained.

The new engine is well-contained behind the Engine vtable with no leakage into common code. The TX and RX free paths correctly separate stride reclaim from metadata-pool return. The CMake invariants (rdma-first ordering, socket always built, DAQIRI_ENGINE_NAME=1 definitions) are preserved. The one outstanding gap is a stale sentence in the configuration walkthrough doc that misses the new MPRQ raw path, which does not affect runtime correctness.

docs/tutorials/configuration-walkthrough.md (line 20 describes ibverbs as roce://-only)

Important Files Changed

Filename Overview
src/engines/ibverbs/daqiri_ibverbs_engine.cpp New 3200-line MPRQ engine: DevX CQ/RQ/TIR, mlx5dv_dr steering, TX WQE-direct path, GPU reorder, send scheduling. BurstParams free contract correctly separated (stride release vs. metadata pool return). TX two-phase handoff (fill-thread → worker → pool) is sound.
include/daqiri/types.h Adds EngineType::IBVERBS and a stream-aware config_engine_from_string overload; updates engine_type_supports_stream_type to accept IBVERBS for RAW streams. Changes are minimal, correct, and properly guarded by DAQIRI_ENGINE_IBVERBS.
src/CMakeLists.txt ibverbs now maps to both 'rdma' and 'ibverbs' internal engines. The rdma-first ordering (for socket→rdma link), socket always-built, and DAQIRI_ENGINE_=1 compile-definition invariants are all preserved.
src/engines/ibverbs/daqiri_ibverbs_engine.h IbverbsEngine declaration with complete Engine vtable override coverage. Several per-queue ibv_wq/ibv_qp/dr_domain fields remain (verbs legacy path) but are harmlessly initialised to nullptr and guarded in shutdown.
docs/tutorials/configuration-walkthrough.md Line 20 still describes ibverbs as enabling only roce:// endpoints, omitting the new MPRQ raw-Ethernet path; needs one-line update per doc-sync rule.

Reviews (6): Last reviewed commit: "Add pure-ibverbs MPRQ raw engine (engine..." | Re-trigger Greptile

Comment on lines +363 to +371
struct ibv_mr* gmr = ibv_reg_dmabuf_mr(pd, offset, mr.ttl_size_, va, dmabuf_fd, access);
if (gmr == nullptr) {
DAQIRI_LOG_CRITICAL("ibv_reg_dmabuf_mr failed for MR {} ({} bytes): {}", mr_name,
mr.ttl_size_, strerror(errno));
close(dmabuf_fd);
return Status::NULL_PTR;
}
*out_base = static_cast<uint8_t*>(base);
*out_lkey = gmr->lkey;

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P1 File descriptor leak on successful GPU MR registration. ibv_reg_dmabuf_mr makes a kernel-side dma_buf_get() call that increases the underlying dma_buf's reference count, so the userspace fd is no longer needed after a successful call. The error path already closes it; the success path does not, leaking one fd per GPU MR per queue for the lifetime of the process.

Suggested change
struct ibv_mr* gmr = ibv_reg_dmabuf_mr(pd, offset, mr.ttl_size_, va, dmabuf_fd, access);
if (gmr == nullptr) {
DAQIRI_LOG_CRITICAL("ibv_reg_dmabuf_mr failed for MR {} ({} bytes): {}", mr_name,
mr.ttl_size_, strerror(errno));
close(dmabuf_fd);
return Status::NULL_PTR;
}
*out_base = static_cast<uint8_t*>(base);
*out_lkey = gmr->lkey;
struct ibv_mr* gmr = ibv_reg_dmabuf_mr(pd, offset, mr.ttl_size_, va, dmabuf_fd, access);
close(dmabuf_fd);
if (gmr == nullptr) {
DAQIRI_LOG_CRITICAL("ibv_reg_dmabuf_mr failed for MR {} ({} bytes): {}", mr_name,
mr.ttl_size_, strerror(errno));
return Status::NULL_PTR;
}
*out_base = static_cast<uint8_t*>(base);
*out_lkey = gmr->lkey;

Comment thread README.md Outdated
Comment on lines +48 to +52
- ibverbs backend: requires a Mellanox/mlx5 NIC (ConnectX-6/7, BlueField).
Multi-queue 5-tuple flow steering is not yet implemented (single-queue flow
matching works via a catch-all rule); HDS is a logical split (headers remain
in the same, possibly GPU, buffer). RX (MPRQ), TX, GPUDirect, and GPU
reorder/quantize are supported.

Copy link
Copy Markdown
Contributor

Choose a reason for hiding this comment

The reason will be displayed to describe this comment to others. Learn more.

P2 The ibverbs limitations bullet was written for the initial milestone commit and not updated as later commits in this PR added multi-queue 5-tuple flow steering (install_port_flows, mlx5dv_dr, per-packet MARK) and physical HDS (Add physical HDS + multi-region RX via a non-striding DevX RQ, Add TX-side HDS). As merged the README will tell operators these features are absent when they are fully present.

Suggested change
- ibverbs backend: requires a Mellanox/mlx5 NIC (ConnectX-6/7, BlueField).
Multi-queue 5-tuple flow steering is not yet implemented (single-queue flow
matching works via a catch-all rule); HDS is a logical split (headers remain
in the same, possibly GPU, buffer). RX (MPRQ), TX, GPUDirect, and GPU
reorder/quantize are supported.
- ibverbs backend: requires a Mellanox/mlx5 NIC (ConnectX-6/7, BlueField).
RX (MPRQ), TX, GPUDirect, multi-queue 5-tuple flow steering (`mlx5dv_dr`,
per-packet MARK), physical header-data split (HDS) on both RX and TX, accurate
TX send scheduling (wait-on-time WAIT WQE), RX hardware timestamps, and GPU
reorder/quantize are supported.

Note: If this suggestion doesn't match your team's coding style, reply to this and let me know. I'll remember it for next time!

@cliffburdick cliffburdick changed the title Add pure-ibverbs MPRQ raw-Ethernet backend (engine: ibverbs) #152: Add pure-ibverbs MPRQ raw-Ethernet backend (engine: ibverbs) Jun 10, 2026
…abstraction

Re-ports the pure-libibverbs/DevX Multi-Packet (striding) Receive Queue
raw-Ethernet engine onto main's new engine abstraction (Engine/EngineFactory,
src/engines/, DAQIRI_ENGINE, stream_type+engine selection). This consolidates
the prior 16-commit feature branch (preserved at backup-ibverbs-premain) onto
the renamed structure introduced by #156.

Engine: src/engines/ibverbs/daqiri_ibverbs_engine.{cpp,h} (IbverbsEngine :
public Engine) + mlx5_prm_min.h. Drives a mlx5 striding RQ via DevX (CQ + RQ +
TIR + mlx5dv_dr steering, manual WQE/doorbell, cyclic refill); RX (MPRQ, host or
GPU via dmabuf), direct-SQ TX, physical/logical HDS, multi-queue 5-tuple flow
steering with per-packet flow IDs, flex-item arbitrary-offset and
IPv4-total-length matching (flex parser / misc_parameters_4), RX hardware
timestamps, accurate TX send scheduling, GPU reorder/quantize, and netdev-MTU
auto-raise for jumbo frames. Reuses DPDK rte_ring/rte_mempool for the
worker->app handoff only.

Selection: new EngineType::IBVERBS. The user-facing engine name "ibverbs" now
resolves stream-type-aware -- "ibverbs"+raw -> IBVERBS (MPRQ), "ibverbs"+socket
-> RDMA (RoCE) -- via a config_engine_from_string(str, stream_type) overload used
by both EngineFactory::get_engine_type and the YAML decode path. raw still
defaults to dpdk. CMake: DAQIRI_ENGINE="...ibverbs" now builds both internal
engines (rdma for RoCE, ibverbs for raw MPRQ); the ibverbs engine links
libibverbs + libmlx5 + DPDK rings.

Docs (README, concepts, getting-started, configuration, AGENTS) updated to
describe ibverbs as a raw-stream engine selectable with engine: "ibverbs".

Validated on ConnectX loopback: engine: ibverbs + raw closed loop (4.3 Mpps
@512B GPUDirect), raw default still selects dpdk, and flex-parser matching is
byte-accurate against the bench's incrementing payload pattern.

Signed-off-by: Cliff Burdick <cburdick@nvidia.com>
@cliffburdick
cliffburdick force-pushed the ibverbs-mprq-backend branch from 5895c3a to 076da59 Compare June 12, 2026 17:43
@awthomp
awthomp merged commit 63a1219 into main Jun 12, 2026
3 checks passed
@cliffburdick
cliffburdick deleted the ibverbs-mprq-backend branch June 16, 2026 17:12
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Labels

None yet

Projects

None yet

Development

Successfully merging this pull request may close these issues.

2 participants