Source code, RTL, benchmarks, and the full thesis on GitHub
The problem
Market makers face a trade-off between adaptability and determinism. Software models, including learned policies, adapt well to changing volatility regimes but run behind the operating system, the network stack, and the memory allocator, so their response time is long and its tail is unpredictable. Fixed hardware pipelines are fast and deterministic but traditionally only run simple, hand-written quoting logic.
Trishul asks whether a volatility-aware policy can be trained in software and then executed directly in the FPGA data path, so the adaptive decision is made at wire speed while the CPU only supervises.
Architecture
The system is split into a deterministic execution layer on the FPGA, which owns the entire tick-to-order critical path, and a C++20 control plane on the CPU, which handles everything that may be slow without touching the critical path.
FPGA critical path (Verilog, Kintex UltraScale+ target)
rx_parserStrips Ethernet, IP, and UDP headers from 10GbE frames at line rate.
itch_decoderFixed-offset extraction of NASDAQ ITCH 5.0 messages.
book_managerL2 book with O(1) best bid and offer registers; order book imbalance features.
strat_decideDSP-accelerated fixed-point systolic array running the quantized policy network.
risk_gateSingle-cycle pre-trade checks: fat-finger, max notional, duplicate orders (SEC 15c3-5 style).
order_encodeEndian-aware NASDAQ OUCH 5.0 packet formatter.
C++20 hybrid control plane (CPU)
- Model weight hot-swap over PCIe AXI-Lite without pausing the FPGA pipeline
- Lock-free SPSC queues (std::atomic, acquire/release) for telemetry
- Pre-allocated memory pools: no malloc or new on the hot path
- Huge-page (MAP_HUGETLB) buffers emulating DPDK/VFIO polling
- Asynchronous logger that keeps disk I/O off the strategy thread
- AVX2-vectorized signal generation and backtest simulation
Key design decisions
Train in floating point, deploy in fixed point
The policy is trained with Proximal Policy Optimization on order book imbalance and VPIN features, then pruned and quantized to 8- and 16-bit fixed point with quantization-aware training so it fits the FPGA's DSP slices and timing budget. The network is small on purpose: every layer costs clock cycles on the critical path.
Risk checks in the data path, not after it
Pre-trade risk runs as a single-cycle gate between the decision and the order encoder. Making risk part of the pipeline, rather than a software check the order has to wait for, keeps the control both mandatory and cheap.
A software stack with no surprises
Where software is unavoidable, the hot path does no dynamic allocation, communicates through single-producer single-consumer queues with explicit acquire/release ordering instead of mutexes, and pushes logging and disk I/O to background threads. The goal is not just a low median but a tight tail.
Stress the system with a generative market
Historical replay rarely contains enough liquidity crises to test a market maker. Trishul includes a synthetic market generator that combines Merton jump-diffusion for price gaps with a Hawkes process (simulated by Ogata thinning) for self-exciting order flow, and a model-selection script that justifies the hybrid against geometric Brownian motion using AIC and BIC.
Scope and limitations
- Trishul is research software. It simulates direct market access protocols and kernel-bypass networking and has never been connected to a live exchange or traded real capital.
- Results come from simulation and RTL testbenches, not from production hardware on an exchange network, so they should be read as design-level evidence rather than production latency.
What I would do next
- Synthesize to a physical card and measure tick-to-order latency with hardware timestamps on the wire.
- Replay real ITCH captures alongside the synthetic generator to check the policy against real order flow.
- Add a deterministic software fallback path with identical semantics for A/B validation against the FPGA.
