Shreyansh Modi

COMPUTER VISION, ML SECURITY RESEARCHER  ·  IIT Roorkee  ·  India

I am an undergraduate student in Electrical Engineering at IIT Roorkee interested in Artificial Intelligence.

My research interests include Computer Vision, Adversarial Machine Learning, and Diffusion Models.I have worked on projects spanning medical image analysis, visual reasoning agents, and cybersecurity in networked control systems.

Recently, I interned as a Machine Learning Engineer at Infigro.ai, where I developed an AI agent for chart and visual-data understanding. Previously, I was a Research Intern at IIT Roorkee, exploring cybersecurity in control systems. I also represented IIT Roorkee at Inter IIT Tech Meet 14.0, where my team secured 4th place.

Off the clock , I enjoy watching movies, exploring new ideas in AI, and diving into interesting research papers.

You can download my resume here ↓.
Shreyansh Modi

Research and Publications

Guidance for Low-Level Perceptual Editing in Unconditional Diffusion Models S. Modi, A. Tomar, and A. Aggarwal.
Generative Models for Computer Vision Workshop, CVPR, 2026.

A method that enables semantic control over unconditional diffusion models (DDPMs) by extracting concept-specific direction vectors from the model's internal h-space (the bottleneck activation of the U-Net architecture) and applying them as targeted perturbations during the reverse generative process.
arXiv · Code
How Many Counterfactuals Does It Take? Probing VLM Hallucinations Through Circuits and Causal Effects A. Gupta, S. Singh, A. Sinha, S. Modi, and A. Tomar.
arXiv preprint, 2026.

A study of the sample complexity of counterfactual robustness for hallucinated outputs in Visual Language Models, introducing a causal influence metric and using circuit discovery to derive empirical bounds on the counterfactual samples needed to reliably detect instability in hallucinated predictions.
arXiv · Code

Projects

VLM Hallucination Mitigation via Self-Verification Decoding An optimization framework I built to mitigate visual hallucinations in multimodal Vision-Language Models, using decoding-time interventions to keep generated text strictly grounded to image inputs. Evaluated on the POPE and CHAIR benchmarks. VLMs · Decoding · POPE
Rethinking-CD A reproducibility study I conducted critiquing contrastive decoding methods for object hallucinations in multimodal LLMs, showing on LLaVA-1.5 (POPE) that the gains come from distribution shifts and plausibility constraints rather than better visual grounding. VLMs · LLaVA · POPE
RE-UltraBreak A reproducibility study auditing safety vulnerabilities in Vision-Language Models, replicating the optimization pipeline to generate universal, transferable adversarial image jailbreaks from a single surrogate model. Adversarial · VLM Security
InVisionDX A deep learning project I developed for medical imaging, capable of identifying diseases like COVID-19, Pneumonia, Tuberculosis, Lung Cancer, and Alzheimer's from scans with over 95% accuracy. PyTorch · CV · Flask
HireSense An open-source NLP tool I built that uses fine-tuned Named Entity Recognition to parse resumes and perfectly match candidates to job descriptions. Python · NER · NLP
ERP API Framework A high-performance backend architecture I designed for ERP management, complete with data analysis modules that pull actionable insights from raw ERP datasets. Python · Flask · JS
Adobe Photo Editor A smart photo editing application our team built for the InterIIT Tech Meet. It uses computer vision to enable context-aware enhancements completely automatically. 4th / all IITs
Reproducibility Study of Diffusion Beats GANs A reproducibility study on OpenAI's foundational paper, where I replicated the unconditional U-Net and noise-robust classifier to implement classifier guidance, trading diversity for high image fidelity to outperform traditional GANs. Diffusion · GANs · U-Net

Contact