Curriculum Coarse-to-Fine Selection for High-IPC Dataset Distillation
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Chen, Yanda, Chen, Gongwei, Zhang, Miao, Guan, Weili, Nie, Liqiang |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Enhancing Diffusion-based Dataset Distillation via Adversary-Guided Curriculum Sampling
von: Zou, Lexiao, et al.
Veröffentlicht: (2025)
von: Zou, Lexiao, et al.
Veröffentlicht: (2025)
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
von: Zhang, Renshan, et al.
Veröffentlicht: (2025)
von: Zhang, Renshan, et al.
Veröffentlicht: (2025)
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
von: Shen, Leyang, et al.
Veröffentlicht: (2024)
von: Shen, Leyang, et al.
Veröffentlicht: (2024)
Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
von: Zhang, Renshan, et al.
Veröffentlicht: (2024)
von: Zhang, Renshan, et al.
Veröffentlicht: (2024)
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)
PersonalAlign: Hierarchical Implicit Intent Alignment for Personalized GUI Agent with Long-Term User-Centric Records
von: Lyu, Yibo, et al.
Veröffentlicht: (2026)
von: Lyu, Yibo, et al.
Veröffentlicht: (2026)
Beyond Quantity: Distribution-Aware Labeling for Visual Grounding
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
von: Zhang, Yichi, et al.
Veröffentlicht: (2025)
HIERAMP: Coarse-to-Fine Autoregressive Amplification for Generative Dataset Distillation
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
von: Zhao, Lin, et al.
Veröffentlicht: (2026)
R^3: Composed Video Retrieval via Reasoning-Guided Recalling and Re-ranking
von: Li, Zixu, et al.
Veröffentlicht: (2026)
von: Li, Zixu, et al.
Veröffentlicht: (2026)
DAgger Diffusion Navigation: DAgger Boosted Diffusion Policy for Vision-Language Navigation
von: Shi, Haoxiang, et al.
Veröffentlicht: (2025)
von: Shi, Haoxiang, et al.
Veröffentlicht: (2025)
Advancing Complex Video Object Segmentation via Tracking-Enhanced Prompt: The 1st Winner for 5th PVUW MOSE Challenge
von: Zhang, Jinrong, et al.
Veröffentlicht: (2026)
von: Zhang, Jinrong, et al.
Veröffentlicht: (2026)
The 1st Winner for 5th PVUW MeViS-Text Challenge: Strong MLLMs Meet SAM3 for Referring Video Object Segmentation
von: He, Xusheng, et al.
Veröffentlicht: (2026)
von: He, Xusheng, et al.
Veröffentlicht: (2026)
Object-Shot Enhanced Grounding Network for Egocentric Video
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
Curriculum Dataset Distillation
von: Ma, Zhiheng, et al.
Veröffentlicht: (2024)
von: Ma, Zhiheng, et al.
Veröffentlicht: (2024)
EgoAdapt: A Multi-Scene Egocentric Adaptation Method for CVPR 2026 HD-EPIC VQA Challenge
von: Chen, Zhiwei, et al.
Veröffentlicht: (2026)
von: Chen, Zhiwei, et al.
Veröffentlicht: (2026)
TempRet: Temporal Enhancement and Two-Stage Reranking for CVPR 2026 EPIC-KITCHENS-100 Multi-Instance Retrieval Challenge
von: Li, Zixu, et al.
Veröffentlicht: (2026)
von: Li, Zixu, et al.
Veröffentlicht: (2026)
OmniEgo-R$^2$: A Routed Reasoning Framework for the 1st Cross-Domain EgoCross Challenge at CVPR 2026
von: Li, Zixu, et al.
Veröffentlicht: (2026)
von: Li, Zixu, et al.
Veröffentlicht: (2026)
EgoAction: Egocentric Action Composition with Reliability-Aware Temporal Fusion for the EPIC-KITCHENS Action Detection Challenge at CVPR 2026
von: Fu, Zhiheng, et al.
Veröffentlicht: (2026)
von: Fu, Zhiheng, et al.
Veröffentlicht: (2026)
Handling Imbalanced Pseudolabels for Vision-Language Models with Concept Alignment and Confusion-Aware Calibrated Margin
von: Wang, Yuchen, et al.
Veröffentlicht: (2025)
von: Wang, Yuchen, et al.
Veröffentlicht: (2025)
StructAlign: Structured Cross-Modal Alignment for Continual Text-to-Video Retrieval
von: Wang, Shaokun, et al.
Veröffentlicht: (2026)
von: Wang, Shaokun, et al.
Veröffentlicht: (2026)
UniCVR: From Alignment to Reranking for Unified Zero-Shot Composed Visual Retrieval
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
von: Wen, Haokun, et al.
Veröffentlicht: (2026)
JFAA: Technical Report for the EPIC-KITCHENS-100 Action Anticipation Challenge at EgoVis 2026
von: Chu, Qiaohui, et al.
Veröffentlicht: (2026)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2026)
HATS: Hardness-Aware Trajectory Synthesis for GUI Agents
von: Shao, Rui, et al.
Veröffentlicht: (2026)
von: Shao, Rui, et al.
Veröffentlicht: (2026)
Uncovering Hidden Connections: Iterative Search and Reasoning for Video-grounded Dialog
von: Zhang, Haoyu, et al.
Veröffentlicht: (2023)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2023)
OSGNet with MLLM Reranking @ Ego4D Episodic Memory Challenge 2026
von: Feng, Yisen, et al.
Veröffentlicht: (2026)
von: Feng, Yisen, et al.
Veröffentlicht: (2026)
LadderMIL: Multiple Instance Learning with Coarse-to-Fine Self-Distillation
von: Wu, Shuyang, et al.
Veröffentlicht: (2025)
von: Wu, Shuyang, et al.
Veröffentlicht: (2025)
HuM-Eval: A Coarse-to-Fine Framework for Human-Centric Video Evaluation
von: Zhang, Bingzi, et al.
Veröffentlicht: (2026)
von: Zhang, Bingzi, et al.
Veröffentlicht: (2026)
Spatial Understanding from Videos: Structured Prompts Meet Simulation Data
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
OSGNet @ Ego4D Episodic Memory Challenge 2025
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
von: Feng, Yisen, et al.
Veröffentlicht: (2025)
Technical Report for Ego4D Long-Term Action Anticipation Challenge 2025
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2025)
HCQA-1.5 @ Ego4D EgoSchema Challenge 2025
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2025)
MARS: Technical Report for the CASTLE Challenge at EgoVis 2026
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
von: Zhang, Haoyu, et al.
Veröffentlicht: (2026)
VISTA: Technical Report for the Ego4D Short-Term Object Interaction Anticipation at EgoVis 2026
von: Chu, Qiaohui, et al.
Veröffentlicht: (2026)
von: Chu, Qiaohui, et al.
Veröffentlicht: (2026)
Detecting Deepfakes via Hamiltonian Dynamics
von: Cheng, Harry, et al.
Veröffentlicht: (2026)
von: Cheng, Harry, et al.
Veröffentlicht: (2026)
Embodied Crowd Counting
von: Long, Runling, et al.
Veröffentlicht: (2025)
von: Long, Runling, et al.
Veröffentlicht: (2025)
Towards Consistent and Efficient Dataset Distillation via Diffusion-Driven Selection
von: Zhong, Xinhao, et al.
Veröffentlicht: (2024)
von: Zhong, Xinhao, et al.
Veröffentlicht: (2024)
FineCIR: Explicit Parsing of Fine-Grained Modification Semantics for Composed Image Retrieval
von: Li, Zixu, et al.
Veröffentlicht: (2025)
von: Li, Zixu, et al.
Veröffentlicht: (2025)
Global Prior Meets Local Consistency: Dual-Memory Augmented Vision-Language-Action Model for Efficient Robotic Manipulation
von: Li, Zaijing, et al.
Veröffentlicht: (2026)
von: Li, Zaijing, et al.
Veröffentlicht: (2026)
DKDM: Data-Free Knowledge Distillation for Diffusion Models with Any Architecture
von: Xiang, Qianlong, et al.
Veröffentlicht: (2024)
von: Xiang, Qianlong, et al.
Veröffentlicht: (2024)
Continuous Knowledge-Preserving Decomposition with Adaptive Layer Selection for Few-Shot Class-Incremental Learning
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
von: Li, Xiaojie, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Enhancing Diffusion-based Dataset Distillation via Adversary-Guided Curriculum Sampling
von: Zou, Lexiao, et al.
Veröffentlicht: (2025) -
FALCON: Resolving Visual Redundancy and Fragmentation in High-resolution Multimodal Large Language Models via Visual Registers
von: Zhang, Renshan, et al.
Veröffentlicht: (2025) -
MoME: Mixture of Multimodal Experts for Generalist Multimodal Large Language Models
von: Shen, Leyang, et al.
Veröffentlicht: (2024) -
Token-level Correlation-guided Compression for Efficient Multimodal Document Understanding
von: Zhang, Renshan, et al.
Veröffentlicht: (2024) -
PUMA: Layer-Pruned Language Model for Efficient Unified Multimodal Retrieval with Modality-Adaptive Learning
von: Lyu, Yibo, et al.
Veröffentlicht: (2025)