The FINN-T framework recently enabled compilation of quantized Transformer models to streaming dataflow FPGA accelerators, demonstrating end-to-end synthesis for signal processing tasks; however, extending this capability to Vision Transformers introduces distinct challenges, including vision-to-sequence conversion compatible with compiler attention patterns and folding constraints in vision classification heads that differ from prior domains. This paper presents FINN-ViT, which addresses these challenges through (1) a preprocessing-based vision-to-sequence interface that preserves ONNX graph compatibility with FINN-T attention transformations, (2) a folding configuration methodology for vision classification under constrained output dimensionality, and (3) a complete end-to-end compilation flow from Brevitas-quantized SimpleViT model to a ZCU104 FPGA bitstream for the ViT accelerator core, with patch extraction and output aggregation remaining on the host. In a single-layer CIFAR-10 case study, the resulting accelerator achieves 606 µs latency at 100 MHz (1,650 images/sec) with 47% LUT utilization, demonstrating that FINN-T compiler transformations and hardware primitives extend to vision inputs without modification.
FINN-ViT: End-to-End Compilation of Vision Transformers to Streaming Dataflow FPGAs
Farooq, Qaisar;Drago, Idilio
2026-01-01
Abstract
The FINN-T framework recently enabled compilation of quantized Transformer models to streaming dataflow FPGA accelerators, demonstrating end-to-end synthesis for signal processing tasks; however, extending this capability to Vision Transformers introduces distinct challenges, including vision-to-sequence conversion compatible with compiler attention patterns and folding constraints in vision classification heads that differ from prior domains. This paper presents FINN-ViT, which addresses these challenges through (1) a preprocessing-based vision-to-sequence interface that preserves ONNX graph compatibility with FINN-T attention transformations, (2) a folding configuration methodology for vision classification under constrained output dimensionality, and (3) a complete end-to-end compilation flow from Brevitas-quantized SimpleViT model to a ZCU104 FPGA bitstream for the ViT accelerator core, with patch extraction and output aggregation remaining on the host. In a single-layer CIFAR-10 case study, the resulting accelerator achieves 606 µs latency at 100 MHz (1,650 images/sec) with 47% LUT utilization, demonstrating that FINN-T compiler transformations and hardware primitives extend to vision inputs without modification.| File | Dimensione | Formato | |
|---|---|---|---|
|
2026_SAS_FINN-ViT_AAM.pdf
Accesso aperto
Descrizione: Author manuscript
Tipo di file:
POSTPRINT (VERSIONE FINALE DELL’AUTORE)
Dimensione
417.32 kB
Formato
Adobe PDF
|
417.32 kB | Adobe PDF | Visualizza/Apri |
I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.



