Ishan Dave

I am a Senior Applied Scientist at Adobe Firefly Foundry in San Jose, California. I work on controllable image and video generation, diffusion models, and vision-language models.

I received my Ph.D. in Computer Science from the Center for Research in Computer Vision at the University of Central Florida in 2024, advised by Prof. Mubarak Shah. My research spans fine-grained video understanding, self-supervised learning, and privacy preservation.

CV (public version, for more details feel free to ask!)

Ishan Dave

Updates

2026
  2026: Selected as Outstanding Reviewer of CVPR 2026! (top 5% among 17,491 reviewers)🥇
  Jun'26: A first author paper "CreativeVR" presented at the J2A Workshop at CVPR 2026
  2026: A paper "Privacy Beyond Pixels" accepted at ICLR 2026
  Jan'26: Our patent on identifying and aligning video clips was granted!
2025
  Dec'25: A First author work "CreativeVR" released: state-of-the-art video restoration model
  Oct'25: The Generative Upscaler project I led shipped as the default Generative Upscale capability in Photoshop 💥
  Sep'25: A paper "Finegrained Video Retrieval" accepted at NeurIPS 2025
  Jun'25: A paper "GT-Loc" accepted at ICCV 2025: Oral presentation! (Top 0.6% of papers)
  Apr'25: A paper "ALBAR" accepted at ICLR 2025
  Jan'25: Started full-time at Adobe Firefly, Seattle, WA
2024
  Oct'24: Successfully defended my Ph.D. dissertation!
  Aug'24: SPAct Patent Approved! my first patent as the primary inventor 💥
  Jul'24: 2 first author papers accepted at ECCV 2024; Sync from the Sea selected for an oral presentation (top 3% of accepted papers)💥💥
  Jun'24: Selected as Outstanding Reviewer of CVPR 2024! (top 2% among 10,000 reviewers)🥇
  May'24: Started internship at Apple, Cupertino, CA
2023
  Dec'23: A First author paper "No More Shortcuts" accepted to AAAI 2024 💥
  Jul'23: A First author paper "Event-TransAct" accepted at IROS 2023 💥
  Jul'23: A paper "TeD-SPAD" accepted at ICCV 2023
  May'23: Started summer internship at Adobe, San Jose, CA
  Mar'23: A First author paper "TimeBalance" accepted to CVPR 2023 💥
  Jan'23: A paper "TransVisDrone" accepted at ICRA 2023
2022
  May'22: Started summer internship at Adobe, USA (remote- Florida)
  Mar'22: A First author paper "TCLR" accepted to CVIU 2022 💥
  Mar'22: A First author paper "SPAct" accepted to CVPR 2022 💥
2021 & Earlier
  Jan'21: Our Gabriella paper has been awarded the best scientific paper award at ICPR 2020

Work Experience

Adobe Firefly FoundrySenior Applied Scientist
Adobe Firefly Foundry
San Jose, CA · January 2025 to present
  • Controllable generation: Lead controllable image and video generation for enterprise workflows, from large-scale pretraining and supervised fine-tuning to custom model adaptation. Shipped multiple 2D and 3D generation controls to leading animation studios, including Disney.
  • Vision-language models: Train and adapt custom VLMs and multimodal LLMs for large-scale data curation and evaluation of image and video generative models.
  • Generative Upscaler: Initiated and led the project from research proposal and prototype through production training and deployment. Shipped as Photoshop’s default Generative Upscale capability, used by 80% of a 15M+ paid user base, and supporting Adobe Firefly Custom Models.
  • Generative Video Refiner / CreativeVR: Initiated and led project direction and model development to restore structural and motion artifacts in generated and real videos. Presented at the J2A Workshop at CVPR 2026.
ApplePh.D. AI/ML Intern
Apple
Cupertino, CA · May to August 2024
  • Enhanced Stable Diffusion for image editing using vision-language and multimodal foundation models.
  • Trained high-resolution diffusion models on a dataset of 10 million samples.
Adobe ResearchResearch Scientist Intern
Adobe Research
San Jose, CA · May to November 2023
Mentors: Dr. Simon Jenni and Dr. Fabian Caba
  • Built and evaluated fine-grained video retrieval over galleries containing millions of videos, extending retrieval beyond semantic similarity to alignable action phases and key events.
  • Resulted in Sync from the Sea, ECCV 2024 Oral (top 3% of accepted papers) and a granted US patent.
Adobe ResearchResearch Scientist Intern
Adobe Research
Remote, USA · May to November 2022
Mentor: Dr. Simon Jenni
  • Developed a self-supervised video representation system with frame-level temporal recognition tasks and augmentations that reduce shortcut learning.
  • Delivered an evaluation suite spanning video classification, retrieval, and temporal correspondence across 10 benchmarks. Published as No More Shortcuts, AAAI 2024.

Publications

My research spans generative image and video models, vision-language models, and fine-grained video understanding. I study controllable generation and restoration, video retrieval and alignment, self-supervised and semi-supervised learning, and privacy-preserving representations. Earlier work includes event-camera action recognition and drone-to-drone detection.

Selected publications are listed from newest to oldest; representative papers are highlighted.

CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos figure
CreativeVR: Diffusion-Prior-Guided Approach for Structure and Motion Restoration in Generative and Real Videos
Tejas Panambur*, Ishan Rajendrakumar Dave*, Chongjian Ge, Ersin Yumer, Xue Bai
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2026
J2A Workshop
*= equal contribution

Modern text-to-video (T2V) diffusion models can synthesize visually compelling clips, yet they remain brittle at fine-scale structure: even state-of-the-art generators often produce distorted faces and hands, warped backgrounds, and temporally inconsistent motion. Such severe structural artifacts also appear in very low-quality real-world videos. Classical video restoration and super-resolution (VR/VSR) methods, in contrast, are tuned for synthetic degradations such as blur and downsampling and tend to stabilize these artifacts rather than repair them, while diffusion-prior restorers are usually trained on photometric noise and offer little control over the trade-off between perceptual quality and fidelity.

We introduce CreativeVR, a diffusion-prior-guided video restoration framework for AI-generated (AIGC) and real videos with severe structural and temporal artifacts. Our deep-adapter-based method exposes a single precision knob that controls how strongly the model follows the input, smoothly trading off between precise restoration on standard degradations and stronger structure- and motion-corrective behavior on challenging content. Our key novelty is a temporally coherent degradation module used during training, which applies carefully designed transformations that produce realistic structural failures.

To evaluate AIGC-artifact restoration, we propose the AIGC54 benchmark with FIQA, semantic and perceptual metrics, and multi-aspect scoring. CreativeVR achieves state-of-the-art results on videos with severe artifacts and performs competitively on standard video restoration benchmarks, while running at practical throughput (~13 FPS @ 720p on a single 80 GB A100).

Privacy Beyond Pixels: latent anonymization overview
Privacy Beyond Pixels: Latent Anonymization for Privacy-Preserving Video Understanding
Joseph Fioresi, Ishan Rajendrakumar Dave, Mubarak Shah
International Conference on Learning Representations (ICLR), 2026

We introduce a lightweight anonymizing adapter that removes private information from the latent features of frozen video foundation models while preserving their usefulness across downstream tasks. Training combines a self-supervised privacy objective, multitask co-training, and latent consistency regularization. The approach reduces privacy leakage while retaining performance on action recognition, temporal action detection, and anomaly detection, and also mitigates gender bias in action recognition.

From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos figure
From Play to Replay: Composed Video Retrieval for Temporally Fine-Grained Videos
Animesh Gupta, Jay Parmar, Ishan Rajendrakumar Dave, Mubarak Shah
Conference on Neural Information Processing Systems (NeurIPS) , 2025

We enable fine-grained video retrieval from a query video and a natural-language modification, focusing on subtle differences in action and timing. TF-CoVR provides 180K triplets from gymnastics and diving videos. Temporally discriminative video representations and contrastive alignment help retrieve clips that satisfy the requested change while preserving the relevant visual context.

GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space figure
GT-Loc: Unifying When and Where in Images Through a Joint Embedding Space
David Shatwell, Ishan Rajendrakumar Dave, Sirnam Swetha, Mubarak Shah
International Conference on Computer Vision (ICCV) , 2025
Oral presentation! (Top 0.6% papers)

We learn a shared embedding space across images, capture time, and geolocation to jointly infer when and where an image was taken. Cyclical temporal metric learning captures the structure of time and supports compositional and text-based image retrieval, with applications in metadata recovery and visual search.

ALBAR: Adversarial Learning approach to mitigate Biases in Action Recognition figure
ALBAR: Adversarial Learning approach to mitigate Biases in Action Recognition
Joseph Fioresi, Ishan Rajendrakumar Dave, Mubarak Shah
International Conference on Learning Representations (ICLR) , 2025
Poster Presentation

We reduce foreground and background shortcut bias in video action recognition through adversarial training on static clips. Entropy maximization and gradient regularization discourage reliance on spurious appearance cues, improving robustness and the learning of action-relevant video representations.

Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets figure
Sync from the Sea: Retrieving Alignable Videos from Large-Scale Datasets
Ishan Rajendrakumar Dave, Fabian Caba, Mubarak Shah, Simon Jenni.
The 18th European Conference on Computer Vision (ECCV) , 2024
Oral presentation! (Top 3% of accepted papers)

We reframe temporal video alignment as a large-scale retrieval problem: find clips that can be synchronized at corresponding action phases and key events. Our approach reranks semantic retrieval results with the DRAQ alignability score, then aligns the best matches using frame-level representations and dynamic time warping. This supports video editing, processing, and understanding workflows.

FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition figure
FinePseudo: Improving Pseudo-Labelling through Temporal-Alignablity for Semi-Supervised Fine-Grained Action Recognition
Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Mubarak Shah.
The 18th European Conference on Computer Vision (ECCV) , 2024

We introduce Alignability-Verification-based Metric learning for semi-supervised fine-grained action recognition. Using dynamic time warping (DTW) for action-phase-aware comparison, our learnable alignability score refines pseudo-labels of the video encoder. Our framework, FinePseudo, outperforms prior methods on fine-grained action recognition datasets. Additionally, it demonstrates robustness in handling novel unlabeled classes in open-world setups.

CodaMal: Contrastive Domain Adaptation for Malaria Detection in Low-Cost Microscopes figure
CodaMal: Contrastive Domain Adaptation for Malaria Detection in Low-Cost Microscopes
Ishan Rajendrakumar Dave, Tristan de Blegiers, Chen Chen, Mubarak Shah.
31st IEEE International Conference on Image Processing (ICIP) , 2024
Oral presentation!

We propose a Domain Adaptive Contrastive objective to bridge the gap between High and Low Cost Microscopes. On the publicly available large-scale M5 dataset, our proposed method shows a significant improvement of 16% over the state-of-the-art methods in terms of the mean average precision metric (mAP), provides a 21× speed-up during inference, and requires only half as many learnable parameters as the prior methods.

No More Shortcuts: Realizing the Potential of Temporal Self-Supervision figure
No More Shortcuts: Realizing the Potential of Temporal Self-Supervision
Ishan Rajendrakumar Dave, Simon Jenni, Mubarak Shah.
AAAI Conference on Artificial Intelligence, Main Technical Track (AAAI) , 2024

We demonstrate experimentally that our more challenging frame-level task formulations and the removal of shortcuts drastically improve the quality of features learned through temporal self-supervision. Our extensive experiments show state-of-the-art performance across 10 video understanding datasets, illustrating the generalization ability and robustness of our learned video representations.

TeD-SPAD: Temporal Distinctiveness for Self-supervised Privacy-preservation for Video Anomaly Detection figure
TeD-SPAD: Temporal Distinctiveness for Self-supervised Privacy-preservation for Video Anomaly Detection
Joseph Fioresi, Ishan Rajendrakumar Dave, Mubarak Shah.
Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV), 2023

We propose TeD-SPAD, a privacy-aware video anomaly detection framework that destroys visual private information in a self-supervised manner. In particular, we propose the use of a temporally-distinct triplet loss to promote temporally discriminative features, which complements current weakly-supervised VAD methods.

EventTransAct: A Video Transformer-based Framework for Event-camera Based Action Recognition figure
EventTransAct: A Video Transformer-based Framework for Event-camera Based Action Recognition
Tristan de Blegiers*, Ishan Rajendrakumar Dave*, Adeel Yousaf, Mubarak Shah.
*= equal contribution
IEEE/RSJ International Conference on Intelligent Robots and Systems (IROS), 2023

We propose a video transformer-based framework for event-camera based action recognition, which leverages event-contrastive loss and augmentations to adapt the network to event data. Our method achieved state-of-the-art results on N-EPIC Kitchens dataset and competitive results on the standard DVS Gesture recognition dataset, while requiring less computation time compared to competitive prior approaches.

TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition figure
TimeBalance: Temporally-Invariant and Temporally-Distinctive Video Representations for Semi-Supervised Action Recognition
Ishan Rajendrakumar Dave, Mamshad Nayeem Rizve, Chen Chen, Mubarak Shah.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2023

We propose a student-teacher semi-supervised learning framework, where we distill knowledge from a temporally-invariant and temporally-distinctive teacher. Depending on the nature of the unlabeled video, we dynamically combine the knowledge of these two teachers based on a novel temporal similarity-based reweighting scheme. State-of-the-art results on Kinetics400, UCF101, HMDB51.

Transvisdrone: Spatio-temporal Transformer for Vision-based Drone-to-drone Detection in Aerial Videos figure
TransVisDrone: Spatio-Temporal Transformer for Vision-Based Drone-to-Drone Detection in Aerial Videos
Tushar Sangam, Ishan Rajendrakumar Dave, Waqas Sultani, Mubarak Shah.
2023 IEEE International Conference on Robotics and Automation (ICRA), 2023

We propose a simple yet effective framework, TransVisDrone, that provides an end-to-end solution with higher computational efficiency. We utilize CSPDarkNet-53 network to learn object-related spatial features and VideoSwin model to improve drone detection in challenging scenarios by learning spatio-temporal dependencies of drone motion.

SPAct: Self-supervised Privacy Preservation for Action Recognition figure
SPAct: Self-supervised Privacy Preservation for Action Recognition
Ishan Rajendrakumar Dave, Chen Chen, Mubarak Shah.
Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR), 2022

For the first time, we present a novel training framework that removes privacy information from input video in a self-supervised manner without requiring privacy labels. We train our framework using a minimax optimization strategy to minimize the action recognition cost function and maximize the privacy cost function through a contrastive self-supervised loss.

TCLR: Temporal Contrastive Learning for Video Representation figure
TCLR: Temporal Contrastive Learning for Video Representation
Ishan Dave, Rohit Gupta, Mamshad Nayeem Rizve, Mubarak Shah.
Computer Vision and Image Understanding (CVIU), 2022
(280+ citations as of August 2026)

We propose a new temporal contrastive learning framework for self-supervised video representation learning, consisting of two novel losses that aim to increase the temporal diversity of learned features. The framework achieves state-of-the-art results on various downstream video understanding tasks, including significant improvement in fine-grained action classification for visually similar classes.

Gabriellav2: Towards Better Generalization in Surveillance Videos for Action Detection figure
GabriellaV2: Towards Better Generalization in Surveillance Videos for Action Detection
Ishan Dave, Zacchaeus Scheffer, Akash Kumar, Sarah Shiraz, Yogesh Singh Rawat, Mubarak Shah.
IEEE/CVF Winter Conference on Applications of Computer Vision (WACV) Workshops, HADCV, 2022

We propose a realtime, online, action detection system which can generalize robustly on any unknown facility surveillance videos. We tackle the challenging nature of action classification problem in various aspects like handling the class-imbalance training using PLM method and learning multi-label action correlations using LSEP loss. In order to improve the computational efficiency of the system, we utilize knowledge distillation.

Gabriella: An Online System for Real-Time Activity Detection in Untrimmed Security Videos figure
Gabriella: An Online System for Real-Time Activity Detection in Untrimmed Security Videos
Mamshad Nayeem Rizve, Ugur Demir, Praveen Tirupattur, Aayush Jung Rana, Kevin Duarte, Ishan R Dave, Yogesh S Rawat, Mubarak Shah.
25th International Conference on Pattern Recognition (ICPR 2020), held in January 2021 (Best Scientific Paper Award)

Gabriella consists of three stages: tubelet extraction, activity classification, and online tubelet merging. Gabriella utilizes a localization network for tubelet extraction, with a novel Patch-Dice loss to handle variations in actor size, and a Tubelet-Merge Action-Split (TMAS) algorithm to detect activities efficiently and robustly.

Patents

Identifying and aligning video clips from large-scale video datasets figure
Identifying and aligning video clips from large-scale video datasets
Simon Jenni, Ishan Rajendrakumar Dave, Fabian Caba
US Patent US12536801B2. (Status: Granted), 2026
Granted January 27, 2026
Self-Supervised Privacy Preservation Action Recognition System figure
Self-Supervised Privacy Preservation Action Recognition System
Ishan Rajendrakumar Dave, Chen Chen, Mubarak Shah
US Patent US12142053B2. (Status: Granted) , 2024
Granted November 12, 2024

Recognition

Professional Service

NSF Research Experiences for Undergraduates

Mentored Kevin Chung (2022), Ethan Thomas (2021), and Kali Carter (2020).