New Infinite Self-Attention accepted at BMVC 2026

Giorgio Roffo PhD

Head of AI & Manager · AI Research Scientist

Self-Attention  /  LLMs  /  Agentic AI  /  Efficient Inference  /  Computer Vision

I turn frontier research into production systems. Today I lead the AI team at Equixly, building LLM-based multi-agent systems for automated API security testing — and I keep publishing on the mathematics of attention.

My research reformulates self-attention as diffusion on a graph of tokens, a line of work that runs from Infinite Feature Selection (ICCV 2015/2017, IEEE T-PAMI) to Infinite Self-Attention (BMVC 2026). Along the way I led the AI R&D behind FDA-cleared medical AI deployed worldwide.

50+ Publications
2,590+ Citations
14+ Years in AI research
23,600+ FSLib downloads
Highlight

Latest research

A new formulation of self-attention, developed across five institutions and accepted at the British Machine Vision Conference.

Accepted BMVC 2026 · Lancaster, UK · November 2026 404 / 1,448 accepted · 27.9%

Self-Attention And Beyond the Infinite: Towards Linear Transformers with Infinite Self-Attention

G. Roffo, H. Abdelkawy, N. Lavie, L. Palmer

Softmax attention costs O(N²), which caps how much context a Transformer can see. We reframe each attention layer as a diffusion step on a content-adaptive graph of tokens, accumulating multi-hop interactions through a discounted Neumann series. This connects attention directly to classical graph centrality — Katz, PageRank, eigenvector centrality — so token importance becomes interpretable rather than a black box. From that foundation we derive Linear-InfSA, an O(N) variant that approximates the principal eigenvector of the attention operator without ever building the N × N matrix — a drop-in replacement for standard Vision Transformer attention.

Equixly API Security · IT Toyota Motor Europe · BE MindVisionLabs · UK UCL Institute of Cognitive Neuroscience · UK GlimpseML · UK
84.7%+3.2 pp ImageNet-1K top-1, vs. 81.5% for an equal-depth softmax ViT
13× Throughput and energy efficiency — 231 img/s at 0.87 J/img on one A100
332K Tokens at inference (9216²), the only tested model to finish without OOM
O(N) Complexity, with a fixed-size auxiliary state independent of sequence length
0.985 Cosine similarity to the dominant eigenvector of the full quadratic operator
Expertise

What I build

The areas I work in hands-on, from the training run to the served endpoint.

LLM Training & Alignment

Fine-tuning and aligning self-hosted models, including safety alignment and security-layer training for LLM agents.

PEFTLoRASFT RLHFDPOGRPO QAT FP8

Efficient Inference

Squeezing latency and cost out of serving: cache strategy, prefill/decode balance, speculative decoding, and low-precision quantization with calibration.

vLLMKV cache EAGLE-3MTP FP8NVFP4

Agentic AI

Tool-calling agents and multi-agent orchestration for autonomous problem solving, including penetration-testing agents that chain actions across services.

RAGReAct ReflexionCodeAct RL + tool use

Attention & Transformers

The research thread: infinite and linear self-attention, graph-theoretic token weighting, and Vision Transformers that scale to high resolution.

InfSALinear attention ViTInf-FS

Leading AI Teams

Running research-to-production under strict ML protocols, plus AI governance: per-feature compliance documentation and risk categorization under the EU AI Act.

ML protocolsData curation EU AI ActRoadmapping

Stack

The tooling I use daily to train, evaluate, and ship models at scale.

PythonPyTorch Hugging FaceDeepSpeed FSDPTensorRT DockerK8s
Career

Where I have worked

Fourteen years across industry AI and academic research. Expand any role for detail.

Jan 2026 — Present

Head of Artificial Intelligence & Manager

Equixly — Agentic API Security·Verona, Italy

Equixly replaces slow, human-led penetration testing with AI agents that test APIs continuously at machine speed inside CI/CD — trusted by leading European banks, insurers, and payment providers. I lead the AI team and direct Equixly AI Research, taking self-hosted LLMs and agentic systems from research into production.

Deployed projects and research
  • Security agents in production: a trained conversational LLM agent with RAG, security layers, and custom-script capability; an agent that repairs OpenAPI/Swagger specifications like a code agent; an exploitation agent with workflow escalation, tool calling, and memory management; LLM-generated penetration-test reports with RAG for business-impact analysis; and an EU AI Act–compliant customer knowledge base.
  • LLM training & fine-tuning: PEFT/LoRA on self-hosted models, SFT and RL alignment, and research on new self-attention mechanisms (infinite and linear attention).
  • Efficient serving with vLLM: KV cache and prefix caching, prefill/decode stage balancing, batching for parallel load, and speculative decoding (EAGLE-3, d-Flash, MTP) with speculated-token tuning.
  • Quantization: FP8 and NVFP4 quantization of open-weight models with calibration, plus Quantization-Aware Training for FP8 models and LoRA adapters.
  • Agentic AI & RL: RL training with tool calling and modern agentic workflows (Reflexion, Self-Refine, Voyager, CodeAct, MetaGPT, ChatDev, Mixture-of-Agents).
  • Security ML pipeline: schema-driven generation, stateful dependency-aware fuzzing, vulnerability-oriented testing, LLM-guided learning, and multi-agent systems — benchmarked on stateful, relationship-aware Docker service endpoints.
  • Governance & management: per-feature compliance documentation and EU AI Act risk categorization; team planning and workload-aware issue assignment with the Product Manager and CTO.

2022 — 2025

Technical Lead & Senior AI Research Scientist

Cosmo Intelligent Medical Devices·Milan, Italy

Led AI R&D for GI Genius, the first FDA-cleared AI system in gastroenterology, deployed worldwide.

Models, data, and delivery
  • Endoscopy vision models: Polyp Sizing, Bowel Preparation Score Estimation, and Polyp Characterization built on ResNet and Vision Transformers via full fine-tuning — FDA-cleared and shipped worldwide on GI Genius.
  • Voice-based clinical reporting: technical lead on two pipelines, a multimodal LLM + TTS system and a standard LLM system, with ASR via Whisper large-v3 and variants.
  • Data curation: DINOv2 embeddings for anomaly detection and clustering, surfacing mislabeled samples to clean the training set.
  • AI governance: led technical due diligence for onboarding third-party partner applications.
  • Full lifecycle: research to prototyping to deployment as regulatory-compliant microservices; published in Medical Image Analysis (2025) and at MICCAI 2024.

2020 — 2022

Principal Investigator / Research Lead

MindVisionLabs (UCL spin-off) & Toyota Motor Europe·London, UK

Autonomous-driving AI research with Prof. Nilli Lavie, head of the UCL Attention and Cognitive Control Lab, in partnership with Toyota Motor Europe's AI Division. This is where the work that became Infinite Self-Attention started.

Research directions
  • Driver attention and distraction detection with Vision Transformers, applying clustering and anomaly detection over embedded video frames.
  • Linear self-attention for large-scale sensor data, improving autonomous-driving perception pipelines.
  • Research output: Self-Attention and Beyond the Infinite — a linear-attention Transformer with infinite self-attention, accepted at BMVC 2026.

2017 — 2020

Senior Research Scientist (Post-Doc)

University of Glasgow·Glasgow, UK

Three-year post-doc funded by an EPSRC investment of £355,564, managing a multi-university research project.

Teaching, mentoring, and collaborations
  • Deep learning and attention models for social signal processing with Prof. Alessandro Vinciarelli; taught one third of the undergraduate Artificial Intelligence course (2019–2020).
  • Mentored MSc and PhD students to graduation on multimodal depression analysis, autism-spectrum video analysis, and attachment disorders, with peer-reviewed publications at IEEE, ACM, and BMVC venues.
  • Established a Glasgow–Stanford collaboration, supported by an appointment as Visiting Postdoctoral Scholar at the Stanford Vision & Learning Lab (Prof. Silvio Savarese, 2019).
  • Weekly collaboration with Heriot-Watt University on the EPSRC-funded SoCoRo project on socially competent robots for autism research — recognized with a SICSA award for excellence in research collaboration.

2014 — 2017

Research Associate / PhD Candidate

University of Verona·Verona, Italy

PhD in Computer Science with Doctor Europaeus certificate. Thesis: Ranking to Learn and Learning to Rank (Prof. Marco Cristani), which has attracted 400+ citations.

Infinite Feature Selection and FSLib
  • Created Infinite Feature Selection (Inf-FS), connecting graph theory to feature selection by formulating selection as a path on a graph — published at ICCV 2015, ICCV 2017, and in IEEE T-PAMI 2020.
  • Authored the Feature Selection Library (FSLib) for MATLAB, 23,600+ downloads, winning the MathWorks Outstanding Contribution Award (2016) and Research Summit invitations in Newton, USA.
  • Awarded the Cooperint International Programme grant, University of Verona (2015).

2012 — 2013

Research Associate

Italian Institute of Technology (IIT)·Genova, Italy

Computer vision and user authentication with Prof. Vittorio Murino (head of PAVIS) and Dr. Loris Bazzani: part-based models for pedestrian detection and tracking, pose estimation, and characterizing human behavior from social-media data. Included a three-month internship at the University of Glasgow School of Psychology classifying expert versus novice CCTV-operator eye movements (CIARP 2013).

Education

Academic background

2020

Post-Doctorate in Computer Vision (Deep Learning)

University of Glasgow — Prof. A. Vinciarelli

2016

PhD in Computer Science, Computer Vision — Doctor Europaeus

University of Verona — Prof. M. Cristani

2013

Master I in Computer Game Development

University of Verona — Prof. U. Castellani

2011

MSc Computer Science, Computer Vision

University of Verona — Prof. M. Cristani

2009

BSc Computer Science, Multimedia

University of Verona — Prof. A. Fusiello

Publications

Selected papers

50+ publications at ICCV, ECCV, IEEE T-PAMI, ACM Multimedia, ACM TOG, ACM CHI, BMVC, and MICCAI.

A Survey of Large Language Models: Foundations and Future Directions

G. Roffo

Preprint, 2025

A Temporal Convolutional Network-Based Approach and a Benchmark Dataset for Colonoscopy Video Temporal Segmentation

C. Biffi, G. Roffo, P. Salvagnini, A. Cherubini

Medical Image Analysis (Elsevier), 2025

Feature Selection Gates with Gradient Routing for Endoscopic Image Computing

G. Roffo, C. Biffi, P. Salvagnini, A. Cherubini

MICCAI 2024 — Springer LNCS

Infinite Feature Selection: A Graph-based Feature Filtering Approach

G. Roffo, S. Melzi, U. Castellani, A. Vinciarelli, M. Cristani

IEEE T-PAMI, 2020

Infinite Latent Feature Selection: A Probabilistic Latent Graph-Based Ranking Approach

G. Roffo, S. Melzi, U. Castellani, A. Vinciarelli

ICCV 2017

Infinite Feature Selection

G. Roffo, S. Melzi, M. Cristani

ICCV 2015

Discrete Time Evolution Process Descriptor for Shape Analysis and Matching

S. Melzi, M. Ovsjanikov, G. Roffo, M. Cristani, U. Castellani

ACM Transactions on Graphics (TOG), 2018

Automating the Administration and Analysis of Psychiatric Tests

G. Roffo, D.-B. Vo, M. Tayarani, et al., A. Vinciarelli

ACM CHI 2019 — Oral

Conversationally-inspired Stylometric Features for Authorship Attribution in Instant Messaging

M. Cristani, G. Roffo, C. Segalin, L. Bazzani, A. Vinciarelli, V. Murino

ACM Multimedia 2012

Recognition

Awards & highlights

Awards & honors

  • 2019 CVPR Outstanding Reviewer Award — IEEE/CVF Conference on Computer Vision and Pattern Recognition.
  • 2019 Rewarding Contribution Award — University of Glasgow.
  • 2018 MathWorks Research Summit — invited participant, Newton, MA, USA.
  • 2017 NVIDIA GPU Research Grant — compute for AI research.
  • 2016 MathWorks Outstanding Contribution Award — for the Feature Selection Library (FSLib).
  • 2015 Cooperint International Programme — University of Verona.
  • — SICSA research-collaboration award — Scottish Informatics and Computer Science Alliance, for excellence in research collaboration.

Highlights

  • 2026 Infinite Self-Attention accepted at BMVC 2026 — 404 of 1,448 submissions accepted, a 27.9% acceptance rate.
  • 2025 FDA-cleared medical AI deployed worldwide on GI Genius, the first FDA-cleared AI system in gastroenterology.
  • 2019 Visiting Postdoctoral Scholar, Stanford Vision & Learning Lab.
  • — EPSRC investment of £355,564 funding a three-year post-doc at the University of Glasgow.
  • — Invited talks at MIT, MathWorks, ICCV, MICCAI, BMVC, and ACM Multimedia.
  • — Languages: Italian (native), English (fluent, C1).
Network

Research collaborations

Working with scholars and engineers across Europe and the United States.

Toyota Motor Europe

AI research on autonomous-driving perception, 2020–2026, with the AI Division in Belgium. Co-author on Infinite Self-Attention.

MindVisionLabs & GlimpseML, UK

UCL spin-off and ML research studio — joint work on attention mechanisms and the BMVC 2026 paper.

University of Verona

Cybersecurity research with the Department of Computer Science (2026), and long-standing work with Prof. Marco Cristani, co-founder of Humatics and Qualyco.

University of Luxembourg

Cybersecurity research on API security and automated testing (2026).

University of Glasgow

Prof. Alessandro Vinciarelli — Director and Principal Investigator at the SOCIAL AI CDT, Advisory Board Member at Substrata.

Stanford University

Visiting Postdoctoral Scholar, Vision & Learning Lab (Prof. Silvio Savarese), 2019.

Get in touch

Let's build something worth publishing

Open to selected research collaborations, invited talks, and advisory work on LLM systems, attention research, and agentic AI.