Grounded multimodal intelligence
We study how models locate visual and textual evidence before they reason—across long documents, tables, figures, and scientific records.
- Vision-language models
- Document intelligence
- Verified reasoning
AIgniteLab · Grounded intelligence
We are a research group exploring multimodal reasoning, scientific search, and structured knowledge for the next generation of foundation models.
Our thesis
Intelligence becomes useful when it is grounded in the right evidence, structured for the problem, and tested against the world.
Research directions
We study how models locate visual and textual evidence before they reason—across long documents, tables, figures, and scientific records.
We build retrieval and agent systems that turn a question into traceable evidence, structured hypotheses, and reproducible research actions.
We use graphs, relational data, and latent structure to make language models more capable, efficient, and resilient to hallucination.
Watch the idea
A short field note on the research questions, systems, and people we hope to bring together.
Selected work

Measuring and reducing data referencing errors in compact language models.

Identifying and leveraging visual evidence retrieval heads in long-context understanding.

Parameter-efficient alignment between graph structure and language model representations.
How we work
A convincing answer is not enough. We want systems that reveal what they used, what they inferred, and where uncertainty remains.
We connect research ideas to real data, evaluation, and usable infrastructure—from scientific retrieval to deployed model pipelines.
Students help define the questions. Early members own complete research arcs and shape the culture of the lab.
People

Principal investigator
Qi works at the intersection of multimodal intelligence, information retrieval, and structured machine learning.
Before starting AIgniteLab, Qi was an Applied Scientist at AWS Bedrock, contributing to large-language-model optimization and systems for structured data, including GraphStorm and GraphRAG. He received his PhD in Computer Science from the University of Illinois Urbana–Champaign, advised by Jiawei Han.

Student collaborator
Rongcan studies the attention mechanisms of vision-language models and how to improve their long-context reasoning.
He is an undergraduate student at Tongji University and has collaborated with Qi since 2025. His work includes VERA, which identifies visual evidence retrieval heads in long-context multimodal understanding.
Join AIgniteLab
We are assembling our founding cohort and welcome research assistants and undergraduate students with a genuine passion for research. The best fit is someone who enjoys turning an unclear, important question into a careful experiment and a working system.
Research tracks
Post-training, verified-reward learning, multimodal data construction, and efficient adaptation of compact language and vision-language models.
PyTorch · model training · data pipelines · RLBuild agents that search scientific literature, use tools, learn from research experience, and assist with reproducible scientific discovery.
LLM agents · retrieval · tool use · research systemsAccelerate model serving and test-time reasoning through dynamic compute allocation, efficient decoding, caching, and systems-level optimization.
Inference systems · CUDA/Triton · serving · optimizationDoctoral researchers are currently hosted through Zhejiang University in Hangzhou. Institutional and degree details are confirmed through the official admissions process.
我们正在招募2027年秋季入学的博士生,也希望招募对研究富有热情的科研 助理和本科生。当前重点方向包括:①小模型与多模态模型训练;②Science Agent与AI for Science系统开发;③模型推理与服务加速。欢迎具有机器 学习、NLP、CV、IR、强化学习或系统背景的同学联系。
Apply by emailSend a CV, transcript, links to your best work, and a short note on a research question you want to pursue.