Case study · MS thesis, Stevens Institute of Technology · Oct 2025 - May 2026

Trishul: An AI-Integrated FPGA Market-Making System

Case study of my Stevens MS thesis: a hardware-software co-designed market maker that puts packet parsing, the order book, reinforcement-learning inference, and pre-trade risk on the FPGA critical path, with a zero-allocation C++20 control plane.

Shreejit Verma

Source code, RTL, benchmarks, and the full thesis on GitHub

The problem

Market makers face a trade-off between adaptability and determinism. Software models, including learned policies, adapt well to changing volatility regimes but run behind the operating system, the network stack, and the memory allocator, so their response time is long and its tail is unpredictable. Fixed hardware pipelines are fast and deterministic but traditionally only run simple, hand-written quoting logic.

Trishul asks whether a volatility-aware policy can be trained in software and then executed directly in the FPGA data path, so the adaptive decision is made at wire speed while the CPU only supervises.

Architecture

The system is split into a deterministic execution layer on the FPGA, which owns the entire tick-to-order critical path, and a C++20 control plane on the CPU, which handles everything that may be slow without touching the critical path.

FPGA critical path (Verilog, Kintex UltraScale+ target)

  1. rx_parser

    Strips Ethernet, IP, and UDP headers from 10GbE frames at line rate.

  2. itch_decoder

    Fixed-offset extraction of NASDAQ ITCH 5.0 messages.

  3. book_manager

    L2 book with O(1) best bid and offer registers; order book imbalance features.

  4. strat_decide

    DSP-accelerated fixed-point systolic array running the quantized policy network.

  5. risk_gate

    Single-cycle pre-trade checks: fat-finger, max notional, duplicate orders (SEC 15c3-5 style).

  6. order_encode

    Endian-aware NASDAQ OUCH 5.0 packet formatter.

C++20 hybrid control plane (CPU)

  • Model weight hot-swap over PCIe AXI-Lite without pausing the FPGA pipeline
  • Lock-free SPSC queues (std::atomic, acquire/release) for telemetry
  • Pre-allocated memory pools: no malloc or new on the hot path
  • Huge-page (MAP_HUGETLB) buffers emulating DPDK/VFIO polling
  • Asynchronous logger that keeps disk I/O off the strategy thread
  • AVX2-vectorized signal generation and backtest simulation

Key design decisions

Train in floating point, deploy in fixed point

The policy is trained with Proximal Policy Optimization on order book imbalance and VPIN features, then pruned and quantized to 8- and 16-bit fixed point with quantization-aware training so it fits the FPGA's DSP slices and timing budget. The network is small on purpose: every layer costs clock cycles on the critical path.

Risk checks in the data path, not after it

Pre-trade risk runs as a single-cycle gate between the decision and the order encoder. Making risk part of the pipeline, rather than a software check the order has to wait for, keeps the control both mandatory and cheap.

A software stack with no surprises

Where software is unavoidable, the hot path does no dynamic allocation, communicates through single-producer single-consumer queues with explicit acquire/release ordering instead of mutexes, and pushes logging and disk I/O to background threads. The goal is not just a low median but a tight tail.

Stress the system with a generative market

Historical replay rarely contains enough liquidity crises to test a market maker. Trishul includes a synthetic market generator that combines Merton jump-diffusion for price gaps with a Hawkes process (simulated by Ogata thinning) for self-exciting order flow, and a model-selection script that justifies the hybrid against geometric Brownian motion using AIC and BIC.

Scope and limitations

What I would do next