The FINN-T framework recently enabled compilation of quantized Transformer models to streaming dataflow FPGA accelerators, demonstrating end-to-end synthesis for signal processing tasks; however, extending this capability to Vision Transformers introduces distinct challenges, including vision-to-sequence conversion compatible with compiler attention patterns and folding constraints in vision classification heads that differ from prior domains. This paper presents FINN-ViT, which addresses these challenges through (1) a preprocessing-based vision-to-sequence interface that preserves ONNX graph compatibility with FINN-T attention transformations, (2) a folding configuration methodology for vision classification under constrained output dimensionality, and (3) a complete end-to-end compilation flow from Brevitas-quantized SimpleViT model to a ZCU104 FPGA bitstream for the ViT accelerator core, with patch extraction and output aggregation remaining on the host. In a single-layer CIFAR-10 case study, the resulting accelerator achieves 606 µs latency at 100 MHz (1,650 images/sec) with 47% LUT utilization, demonstrating that FINN-T compiler transformations and hardware primitives extend to vision inputs without modification.

FINN-ViT: End-to-End Compilation of Vision Transformers to Streaming Dataflow FPGAs

Farooq, Qaisar;Drago, Idilio
2026-01-01

Abstract

The FINN-T framework recently enabled compilation of quantized Transformer models to streaming dataflow FPGA accelerators, demonstrating end-to-end synthesis for signal processing tasks; however, extending this capability to Vision Transformers introduces distinct challenges, including vision-to-sequence conversion compatible with compiler attention patterns and folding constraints in vision classification heads that differ from prior domains. This paper presents FINN-ViT, which addresses these challenges through (1) a preprocessing-based vision-to-sequence interface that preserves ONNX graph compatibility with FINN-T attention transformations, (2) a folding configuration methodology for vision classification under constrained output dimensionality, and (3) a complete end-to-end compilation flow from Brevitas-quantized SimpleViT model to a ZCU104 FPGA bitstream for the ViT accelerator core, with patch extraction and output aggregation remaining on the host. In a single-layer CIFAR-10 case study, the resulting accelerator achieves 606 µs latency at 100 MHz (1,650 images/sec) with 47% LUT utilization, demonstrating that FINN-T compiler transformations and hardware primitives extend to vision inputs without modification.
2026
2026 IEEE Sensors Applications Symposium (SAS)
Vitória, Brazil
15-17 July 2026
2026 IEEE Sensors Applications Symposium (SAS)
IEEE
1
6
FPGA; Vision Transformer; Hardware Acceleration; FINN; Streaming Dataflow; Edge Computing; Real-time Processing
Farooq, Qaisar; Drago, Idilio
File in questo prodotto:
File Dimensione Formato  
2026_SAS_FINN-ViT_AAM.pdf

Accesso aperto

Descrizione: Author manuscript
Tipo di file: POSTPRINT (VERSIONE FINALE DELL’AUTORE)
Dimensione 417.32 kB
Formato Adobe PDF
417.32 kB Adobe PDF Visualizza/Apri

I documenti in IRIS sono protetti da copyright e tutti i diritti sono riservati, salvo diversa indicazione.

Utilizza questo identificativo per citare o creare un link a questo documento: https://hdl.handle.net/2318/2162895
Citazioni
  • ???jsp.display-item.citation.pmc??? ND
  • Scopus ND
  • ???jsp.display-item.citation.isi??? ND
social impact