Publications
273 papers, newest first.
2026
-
Camera Control for Text-to-Image Generation via Learning Viewpoint Tokens
arXiv:2604.19954, 2026. preprint
-
Compositional Reasoning via Joint Image and Language Decomposition
EACL (Findings), 5753-5775, 2026.
-
CRAFT: A Tendon-Driven Hand with Hybrid Hard-Soft Compliance
arXiv:2603.12120, 2026. preprint
-
DeVA: Decoupled Video-Action Model with physical guidance for robot policy learning
arXiv:2607.24159, 2026. preprint
-
EgoVerse: An Egocentric Human Dataset for Robot Learning from Around the World
arXiv:2604.07607, 2026. preprint
-
Generalizable VLA Finetuning via Representation Anchoring and Language-Action Alignment
arXiv:2607.13429, 2026. preprint
-
Resolving Interference (RI): Disentangling Models for Improved Model Merging
arXiv:2603.13467, 2026. preprint
-
Unified Visuomotor Targets: Supervising VLAs Beyond Physical Actions
arXiv:2608.03563, 2026. preprint
-
VISOR: VIsual Spatial Object Reasoning for Language-driven Object Navigation
arXiv:2602.07555, 2026. preprint
2025
-
Constructive Distortion: Improving MLLMs with Attention-Guided Image Warping
arXiv:2510.09741, 2025. preprint
-
GeoDiff: Geometry-Guided Diffusion for Metric Depth Estimation
ICCVW, 6357-6367, 2025.
-
GViT: Representing Images as Gaussians for Visual Recognition
arXiv:2506.23532, 2025. preprint
-
Ouroboros: Single-Step Diffusion Models for Cycle-Consistent Forward and Inverse Rendering
ICCV, 10386-10397, 2025.
-
Robotic Manipulation by Imitating Generated Videos Without Physical Demonstrations
arXiv:2507.00990, 2025. preprint
-
MICCAI (4), 695-705, 2025.
2024
-
AFTer-SAM: Adapting SAM with Axial Fusion Transformer for Medical Imaging Segmentation
WACV, 7960-7969, 2024.
-
CriSp: Leveraging Tread Depth Maps for Enhanced Crime-Scene Shoeprint Matching
ECCV (63), 217-235, 2024.
-
CVTHead: One-shot Controllable Head Avatar with Vertex-feature Transformer
WACV, 6119-6129, 2024.
-
Joint Depth Prediction and Semantic Segmentation with Multi-View SAM
WACV, 1317-1327, 2024.
-
LidaRF: Delving into Lidar for Neural Radiance Field on Street Scenes
CVPR, 19563-19572, 2024.
-
Make the Pertinent Salient: Task-Relevant Reconstruction for Visual Control with Distractions
arXiv:2410.09972, 2024. preprint
-
MaskINT: Video Editing via Interpolative Non-autoregressive Masked Transformers
CVPR, 7403-7412, 2024.
2023
-
Creating a Forensic Database of Shoeprints from Online Shoe-Tread Photos
WACV, 858-868, 2023.
-
GeoFill: Reference-Based Image Inpainting with Better Geometric Understanding
WACV, 1776-1786, 2023.
-
Localized Region Contrast for Enhancing Self-supervised Learning in Medical Image Segmentation
MICCAI (2), 468-478, 2023.
-
MedGen3D: A Deep Generative Framework for Paired 3D Image and Mask Generation
MICCAI (1), 759-769, 2023.
-
Representation Recovering for Self-Supervised Pre-training on Medical Images
WACV, 2684-2694, 2023.
2022
-
AFTer-UNet: Axial Fusion Transformer UNet for Medical Image Segmentation
WACV, 3270-3280, 2022.
-
Automated identification of diverse Neotropical pollen samples using convolutional neural networks
Methods in Ecology and Evolution, 13, 2049--2064, 2022.
-
EI-CLIP: Entity-aware Interventional Contrastive Learning for E-commerce Cross-modal Retrieval
CVPR, 18030-18040, 2022.
-
Geometric Pose Affordance: Monocular 3D Human Pose Estimation with Scene Constraints
ECCV Workshops (6), 3-18, 2022.
-
Identity-Aware Hand Mesh Estimation and Personalization from RGB Images
ECCV (5), 536-553, 2022.
-
PPT: Token-Pruned Pose Transformer for Monocular and Multi-view Human Pose Estimation
ECCV (5), 424-442, 2022.
-
CVPR Workshops, 2317-2326, 2022.
-
Topology-Preserving Shape Reconstruction and Registration via Neural Diffeomorphic Flow
CVPR, 20813-20823, 2022.
2021
-
A linearized framework and a new benchmark for model selection for fine-tuning
arXiv:2102.00084, 2021. preprint
-
Research Square, 2021. preprint
-
Camera Pose Matters: Improving Depth Prediction by Mitigating Pose Distribution Bias
CVPR, 15759-15768, 2021.
-
Recurrent Mask Refinement for Few-Shot Medical Image Segmentation
ICCV, 3898-3908, 2021.
-
Representation Consolidation for Training Expert Students
arXiv:2107.08039, 2021. preprint
-
Scalable and Stable Surrogates for Flexible Classifiers with Fairness Constraints
NeurIPS, 34, 30023-30036, 2021.
-
Sparse Representations for Object- and Ego-Motion Estimations in Dynamic Scenes
IEEE Trans. Neural Networks Learn. Syst., 32, 2521-2534, 2021.
-
Spatial Context-Aware Self-Attention Model For Multi-Organ Segmentation
WACV, 938-948, 2021.
-
Temporal-Aware Self-Supervised Learning for 3D Hand Pose and Mesh Estimation in Videos
WACV, 1049-1058, 2021.
-
Test-Time Training for Deformable Multi-Scale Image Registration
ICRA, 13618-13625, 2021.
2020
-
Celeganser: Automated Analysis of Nematode Morphology and Age
CVPR Workshops, 4164-4173, 2020.
-
Clouds of Oriented Gradients for 3D Detection of Objects, Surfaces, and Indoor Scene Layouts
IEEE Trans. Pattern Anal. Mach. Intell., 42, 2670-2683, 2020.
-
Fine-grained facial expression analysis using dimensional emotion model
Neurocomputing, 392, 38-49, 2020.
-
Proceedings of the National Academy of Sciences, 117, 28496--28505, 2020.
-
Neurogastroenterology & Motility, 33, e14014--e14014, 2020.
-
Nonparametric Structure Regularization Machine for 2D Hand Pose Estimation
WACV, 2020, 370-379, 2020.
-
Predicting Camera Viewpoint Improves Cross-Dataset Generalization for 3D Human Pose Estimation
ECCV Workshops (2), 523-540, 2020.
-
Resisting Large Data Variations via Introspective Transformation Network
WACV, 3069-3078, 2020.
-
Rotation-invariant Mixed Graphical Model Network for 2D Hand Pose Estimation
WACV, 1535-1544, 2020.
-
Weak Supervision and Referring Attention for Temporal-Textual Association Learning
arXiv:2006.11747, 2020. preprint
2019
-
3D Scene Reconstruction With Multi-Layer Depth and Epipolar Transformers
ICCV, 2172-2182, 2019.
-
A Combined Deep Learning-Gradient Boosting Machine Framework for Fluid Intelligence Prediction
ABCD-NP@MICCAI, 1-8, 2019.
-
CeMNet: Self-Supervised Learning for Accurate Continuous Ego-Motion Estimation
CVPR Workshops, 354-363, 2019.
-
Cross-Domain Image Matching with Deep Feature Maps
Int. J. Comput. Vis., 127, 1738-1750, 2019.
-
Nature Communications, 10, 1944--1944, 2019.
-
Journal of Neuroscience, 40, 585--604, 2019.
-
Multi-layer Depth and Epipolar Feature Transformers for 3D Scene Reconstruction
CVPR Workshops, 39-43, 2019.
-
Multigrid Predictive Filter Flow for Unsupervised Learning on Videos
arXiv:1904.01693, 2019. preprint
-
NoduleNet: Decoupled False Positive Reduction for Pulmonary Nodule Detection and Segmentation
MICCAI (6), 266-274, 2019.
-
G3 Genes Genomes Genetics, 9, 2171--2182, 2019.
-
VTNFP: An Image-Based Virtual Try-On Network With Body and Clothing Feature Preservation
ICCV, 10510-10519, 2019.
-
Weakly-Supervised Action Localization With Background Modeling
ICCV, 5501-5510, 2019.
2018
-
A Simple and Effective Fusion Approach for Multi-frame Optical Flow Estimation
ECCV Workshops (6), 706-710, 2018.
-
Active Testing: An Efficient and Robust Framework for Estimating Accuracy
ICML, 3756-3765, 2018.
-
Adversarial deep structured nets for mass segmentation from mammograms
ISBI, 847-850, 2018.
-
bioRxiv (Cold Spring Harbor Laboratory), 2018. preprint
-
DeepEM: Deep 3D ConvNets with EM for Weakly Supervised Pulmonary Nodule Detection
MICCAI (2), 812-820, 2018.
-
-
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 3061-3069, 2018.
-
Recurrent Pixel Embedding for Instance Grouping
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 9018-9028, 2018.
-
Recurrent Scene Parsing with Perspective Understanding in the Loop
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 956-965, 2018.
-
2017
-
Cluster-wise Ratio Tests for Fast Camera Localization
IEEE Conference on Computer Vision and Pattern Recognition, Workshop on Visual Odometry and Computer Vision Applications Base don Location Clues, 2017.
-
-
-
Learning Optimal Parameters for Multi-target Tracking with Contextual Interactions
International Journal of Computer Vision, 122, 484--501, 2017.
-
-
Low-rank Bilinear Pooling for Fine-grained Classification
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2017.
-
Space-Time Localization and Mapping
IEEE International Conference on Computer Vision, 2017.
-
Tracking Objects with Higher Order Interactions using Delayed Column Generation
IEEE Conference on Artificial Intelligence and Statistics, 2017.
2016
-
-
-
-
-
-
Spatially Aware Dictionary Learning and Coding for Fossil Pollen Identification
CVPR workshop CVMI, 2016.
-
-
-
2015
-
Development, 142, 587--596, 2015.
-
BMC Bioinformatics, 16:397, 2015.
-
Articulated Pose Estimation With Tiny Synthetic Videos
The IEEE Conference on Computer Vision and Pattern Recognition (CVPR) Workshops, 2015.
-
Depth-based hand pose estimation: data, methods, and challenges
IEEE International Conference on Computer Vision, 2015.
-
-
-
Hierarchical Planar Correlation Clustering for Cell Segmentation
Energy Minimization Methods in Computer Vision and Pattern Recognition, 492--504, 2015.
-
Learning Optimal Parameters for Multi-target Tracking
British Machine Vision Conference (BMVC), 2015.
-
IEEE International Conference on Computer Vision, 2015.
-
Multi-scale recognition with DAG-CNNs
IEEE International Conference on Computer Vision, 2015.
-
-
Planar Ultrametrics for Image Segmentation
Neural Information Processing Systems (NIPS), 2015.
-
Understanding Everyday Hands in Action from RGB-D Images
IEEE International Conference on Computer Vision, 2015.
-
Using segmentation to predict the absence of occluded parts
British Machine Vision Conference (BMVC), 2015.
-
Nature Protocols, 10, 1860--1896, 2015.
2014
-
3D Hand Pose Detection in Egocentric RGB-D Images
arXiv:1412.0065, 2014. preprint
-
Analysis by synthesis: 3d object recognition by object reconstruction
Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on, 2449--2456, 2014.
-
Capturing long-tail distributions of object subcategories
Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on, 915--922, 2014.
-
-
Learning Multi-target Tracking with Quadratic Object Interactions
arXiv:1412.2066, 2014. preprint
-
-
-
-
Parsing videos of actions with segmental grammars
Computer Vision and Pattern Recognition (CVPR), 2014 IEEE Conference on, 612--619, 2014.
2013
-
Accurate Motion Deblurring using Camera Motion Tracking and Scene Depth
IEEE Workshop on Applications of Computer Vision (WACV), 2013.
-
-
-
International Journal of Computer Vision, 101, 184-204, 2013.
-
-
Monocular 3-D Gait Tracking in Surveillance Scenes
Cybernetics, IEEE Transactions on, PP, 2013.
-
Self-paced learning for long-term tracking
Computer Vision and Pattern Recognition (CVPR), 2013 IEEE Conference on, 2379--2386, 2013.
2012
-
Analyzing 3D Objects in Cluttered Images
Advances in Neural Information Processing Systems 25, 602--610, 2012.
-
Detecting Actions, Poses, and Objects with Relational Phraselets
ECCV (4), 158-172, 2012.
-
Detecting Activities of Daily Living in First-person Camera Views
Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, 2012.
-
Discriminative Decorrelation for Clustering and Classification
ECCV (4), 459-472, 2012.
-
Do we need more training data or better models for object detection?
British Machine Vision Conference (BMVC), 2012.
-
Face detection, pose estimation and landmark estimation in the wild
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2012.
-
Fast Human Pose Detection Using Randomized Hierarchical Cascades of Rejectors
International Journal of Computer Vision, 99, 25-52, 2012.
-
-
-
Patch Mosaic for Fast Motion Deblurring
Asian Conference on Computer Vision (ACCV), 2012.
-
ACS Chemical Neuroscience, 3, 433--438, 2012.
-
-
Computer Vision and Pattern Recognition (CVPR), 2012 IEEE Conference on, 2012.
2011
-
PLoS Genetics, 7, e1002346, 2011.
-
-
Analysis of Gap Gene Reguation in a 3D Organism-Scale Model of the Drosophila melanogaster Embryo
PLoS ONE, 6, e26797, 2011.
-
-
-
Discriminative models for multi-class object layout
International Journal of Computer Vision, 2011.
-
Globally-Optimal Greedy Algorithms for Tracking a Variable Number of Objects
IEEE conference on Computer Vision and Pattern Recognition (CVPR), 2011.
-
Proc. of the International Confernece on Computer Vision, 2011.
-
-
Local Distance Functions: A Taxonomy, New Algorithms, and an Evaluation
IEEE Transactions on Pattern Analysis and Machine Intelligence (PAMI), 2011.
-
Tissue Engineering: Part C, 17(5), 2011.
-
-
-
-
Advanced Materials, 23, 5785-5791, 2011.
-
-
-
-
Where's Waldo: Matching People in Images of Crowds
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2011.
2010
-
-
Discriminative models for static human-object interactions
IEEE Conference on Computer Vision and Pattern Recognition, Workshop on Structured Prediction, 2010.
-
Efficiently Scaling Up Video Annotation with Crowdsourced Marketplaces
Proc. of the European Conference on Computer Vision, 2010.
-
Integrating Data Clustering and Visualization for the Analysis of 3D Gene Expression Data
IEEE Transactions on Computational Biology and Bioinformatics, 7(1), 64-79, 2010.
-
-
-
Jounral of Neuroscience, 30(21), 7269-7280, 2010.
-
BMC Bioinformatics, 11:413, 2010.
-
Robust Tracking of the Upper Limb for Functional Stroke Assessment
IEEE Transactions on Neural Systems and Rehabilitation Engineering (NSRE), 2010.
-
Journal of Real-Time Image Processing (JRTIP), 2010.
2009
-
-
Discriminative models for multi-class object layout
IEEE International Conference on Computer Vision, 2009.
-
Multimedia and Computer Networks (MMCN), 2009.
-
-
Local Distance Functions: A Taxonomy, New Algorithms, and an Evaluation
International Conference on Computer Vision (ICCV), 2009.
-
Object Detection with Discriminatively Trained Part-Based Models
IEEE Pattern Analysis and Machine Intelligence (PAMI), 2009.
-
IEEE Transactions on Computational Biology and Bioinformatics, 6(2), 296-309, 2009.
2008
-
A Discriminatively Trained, Multiscale, Deformable Part Model
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2008.
-
A Quantitative Spatiotemporal Atlas of Gene Expression in the Drosohpila Blastoderm
Cell, 133, 364-374, 2008.
-
Increasing the density of active appearance models
IEEE Conference on Computer Vision and Pattern Recognition (CVPR), 2008.
-
-
-
2007
-
Assessment of Post Stroke Functioning Using Machine Vision
IAPR Conference on Machine Vision Applications (MVA), 2007.
-
-
Learning to parse images of articulated bodies
Advances in Neural Information Processing Systems 19, 1129--1136, 2007.
-
-
Local figure-ground cues are valid for natural images
Journal of Vision, 7, 1--9, 2007.
-
-
2006
-
Genome Biology, 7:R123, 2006.
-
3D Morphology and Gene Expression in the Drosophila Blastoderm at Cellular Resolution II: Dynamics
Genome Biology, 7:R124, 2006.
-
-
Cue Integration for Figure/Ground Labeling
Advances in Neural Information Processing Systems 18, 1121--1128, 2006.
-
-
-
Eurographics/IEEE VGTC Symposium on Visualization, 203--210, 2006.
-
The Rate Adapting Poisson Model for Information Retrieval and Object Recognition
Proceedings of the 23rd International Conference on Machine Learning (ICML 2006), 337-344, 2006.
-
Topographic Product Models Applied to Natural Scene Statistics
Neural Comput., 18, 381--414, 2006.
-
2005
-
Combining Generative Models and Fisher Kernels for Object Recognition
ICCV, I: 136-143, 2005.
-
Computational studies of human motion: part 1, tracking and motion synthesis
Found. Trends. Comput. Graph. Vis., 1, 77--254, 2005.
-
Detecting, Localizing and Recovering Kinematics of Textured Animals
CVPR, II: 635-642, 2005.
-
-
CSB 2005 Workshop on BioImage Data Minning and Informatics, 2005.
-
Scale-Invariant Contour Completion Using Conditional Random Fields
ICCV, II: 1214-1221, 2005.
-
-
2004
-
Automatic Annotation of Everyday Movements
Advances in Neural Information Processing Systems 16, 2004.
-
-
Learning to detect natural image boundaries using local brightness, color, and texture cues
IEEE PAMI, 26, 530-549, 2004.
-
2003
2002
-
Extracting Global Structure from Gene Expression Profiles
Methods of Microarray Data Analysis II, 2002.
-
Learning to Detect Natural Image Boundaries Using Brightness and Texture
Advances in Neural Information Processing Systems, 2002.
-
Spectral Partitioning with Indefinite Kernels Using the Nyström Extension
ECCV, III: 531 ff., 2002.
2001
2000
-
Nonlinear Image Interpolation Through Extended Permutation Filters
ICIP, Vol I: 912-915, 2000.
-
-
-