Enregistré dans:
| Auteurs principaux: | Li, Pengyi, Abdullaeva, Irina, Gambashidze, Alexander, Kuznetsov, Andrey, Oseledets, Ivan |
|---|---|
| Format: | Preprint |
| Publié: |
2025
|
| Sujets: | |
| Accès en ligne: | https://arxiv.org/abs/2502.03183 |
| Tags: |
Ajouter un tag
Pas de tags, Soyez le premier à ajouter un tag!
|
Documents similaires
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
par: Gambashidze, Alexander, et autres
Publié: (2025)
par: Gambashidze, Alexander, et autres
Publié: (2025)
Listener-Rewarded Thinking in VLMs for Image Preferences
par: Gambashidze, Alexander, et autres
Publié: (2025)
par: Gambashidze, Alexander, et autres
Publié: (2025)
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
par: Li, Pengyi, et autres
Publié: (2026)
par: Li, Pengyi, et autres
Publié: (2026)
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
par: Li, Pengyi, et autres
Publié: (2025)
par: Li, Pengyi, et autres
Publié: (2025)
OmniFusion Technical Report
par: Goncharova, Elizaveta, et autres
Publié: (2024)
par: Goncharova, Elizaveta, et autres
Publié: (2024)
From Frames to Clips: Training-free Adaptive Key Clip Selection for Long-Form Video Understanding
par: Sun, Guangyu, et autres
Publié: (2025)
par: Sun, Guangyu, et autres
Publié: (2025)
CoMa: Contextual Massing Generation with Vision-Language Models
par: Maslov, Evgenii, et autres
Publié: (2026)
par: Maslov, Evgenii, et autres
Publié: (2026)
MindShift: Analyzing Language Models' Reactions to Psychological Prompts
par: Vasiliuk, Anton, et autres
Publié: (2025)
par: Vasiliuk, Anton, et autres
Publié: (2025)
Tensor-Train Point Cloud Compression and Efficient Approximate Nearest-Neighbor Search
par: Novikov, Georgii, et autres
Publié: (2024)
par: Novikov, Georgii, et autres
Publié: (2024)
Enhancing Train-Free Infinite-Frame Generation for Consistent Long Videos
par: Feng, X., et autres
Publié: (2026)
par: Feng, X., et autres
Publié: (2026)
Training-Free Action Recognition and Goal Inference with Dynamic Frame Selection
par: Keat, Ee Yeo, et autres
Publié: (2024)
par: Keat, Ee Yeo, et autres
Publié: (2024)
Aligning Diffusion Models with Noise-Conditioned Perception
par: Gambashidze, Alexander, et autres
Publié: (2024)
par: Gambashidze, Alexander, et autres
Publié: (2024)
KTV: Keyframes and Key Tokens Selection for Efficient Training-Free Video LLMs
par: Song, Baiyang, et autres
Publié: (2026)
par: Song, Baiyang, et autres
Publié: (2026)
NoReGeo: Non-Reasoning Geometry Benchmark
par: Abdullaeva, Irina, et autres
Publié: (2026)
par: Abdullaeva, Irina, et autres
Publié: (2026)
Efficient Distribution Matching of Representations via Noise-Injected Deep InfoMax
par: Butakov, Ivan, et autres
Publié: (2024)
par: Butakov, Ivan, et autres
Publié: (2024)
ESQA: Event Sequences Question Answering
par: Abdullaeva, Irina, et autres
Publié: (2024)
par: Abdullaeva, Irina, et autres
Publié: (2024)
M-LLM Based Video Frame Selection for Efficient Video Understanding
par: Hu, Kai, et autres
Publié: (2025)
par: Hu, Kai, et autres
Publié: (2025)
Event-Anchored Frame Selection for Effective Long-Video Understanding
par: Chen, Wang, et autres
Publié: (2026)
par: Chen, Wang, et autres
Publié: (2026)
Adaptive Greedy Frame Selection for Long Video Understanding
par: Huang, Yuning, et autres
Publié: (2026)
par: Huang, Yuning, et autres
Publié: (2026)
Weak-to-Strong 3D Object Detection with X-Ray Distillation
par: Gambashidze, Alexander, et autres
Publié: (2024)
par: Gambashidze, Alexander, et autres
Publié: (2024)
Frame Guidance: Training-Free Guidance for Frame-Level Control in Video Diffusion Models
par: Jang, Sangwon, et autres
Publié: (2025)
par: Jang, Sangwon, et autres
Publié: (2025)
KFS-Bench: Comprehensive Evaluation of Key Frame Sampling in Long Video Understanding
par: Li, Zongyao, et autres
Publié: (2025)
par: Li, Zongyao, et autres
Publié: (2025)
From Captions to Keyframes: KeyScore for Multimodal Frame Scoring and Video-Language Understanding
par: Lin, Shih-Yao, et autres
Publié: (2025)
par: Lin, Shih-Yao, et autres
Publié: (2025)
Speech-to-LaTeX: New Models and Datasets for Converting Spoken Equations and Sentences
par: Korzh, Dmitrii, et autres
Publié: (2025)
par: Korzh, Dmitrii, et autres
Publié: (2025)
Think-Clip-Sample: Slow-Fast Frame Selection for Video Understanding
par: Tan, Wenhui, et autres
Publié: (2026)
par: Tan, Wenhui, et autres
Publié: (2026)
DynImg: Key Frames with Visual Prompts are Good Representation for Multi-Modal Video Understanding
par: Bao, Xiaoyi, et autres
Publié: (2025)
par: Bao, Xiaoyi, et autres
Publié: (2025)
Spread them Apart: Towards Robust Watermarking of Generated Content
par: Pautov, Mikhail, et autres
Publié: (2025)
par: Pautov, Mikhail, et autres
Publié: (2025)
Wavelet-based Frame Selection by Detecting Semantic Boundary for Long Video Understanding
par: Chen, Wang, et autres
Publié: (2026)
par: Chen, Wang, et autres
Publié: (2026)
DiffuseSlide: Training-Free High Frame Rate Video Generation Diffusion
par: Hwang, Geunmin, et autres
Publié: (2025)
par: Hwang, Geunmin, et autres
Publié: (2025)
Frame by Familiar Frame: Understanding Replication in Video Diffusion Models
par: Rahman, Aimon, et autres
Publié: (2024)
par: Rahman, Aimon, et autres
Publié: (2024)
FRAG: Frame Selection Augmented Generation for Long Video and Long Document Understanding
par: Huang, De-An, et autres
Publié: (2025)
par: Huang, De-An, et autres
Publié: (2025)
Geological Field Restoration through the Lens of Image Inpainting
par: Trifonov, Vladislav, et autres
Publié: (2025)
par: Trifonov, Vladislav, et autres
Publié: (2025)
OCC-RAG: Optimal Cognitive Core for Faithful Question Answering
par: Savkin, Maksim, et autres
Publié: (2026)
par: Savkin, Maksim, et autres
Publié: (2026)
A case study of spatiotemporal forecasting techniques for weather forecasting
par: Sofi, Shakir Showkat, et autres
Publié: (2022)
par: Sofi, Shakir Showkat, et autres
Publié: (2022)
Shot-Aware Frame Sampling for Video Understanding
par: Zhao, Mengyu, et autres
Publié: (2026)
par: Zhao, Mengyu, et autres
Publié: (2026)
Latent Inter-Frame Pruning: A Training-Free Method Bridging Traditional Video Compression and Modern Diffusion Transformers for Efficient Generation
par: Menn, Dennis, et autres
Publié: (2026)
par: Menn, Dennis, et autres
Publié: (2026)
One Token per Highly Selective Frame: Towards Extreme Compression for Long Video Understanding
par: Zhang, Zheyu, et autres
Publié: (2026)
par: Zhang, Zheyu, et autres
Publié: (2026)
Graph-to-Frame RAG: Visual-Space Knowledge Fusion for Training-Free and Auditable Video Reasoning
par: Yang, Songyuan, et autres
Publié: (2026)
par: Yang, Songyuan, et autres
Publié: (2026)
Generative Frame Sampler for Long Video Understanding
par: Yao, Linli, et autres
Publié: (2025)
par: Yao, Linli, et autres
Publié: (2025)
Logit-KL Flow Matching: Non-Autoregressive Text Generation via Sampling-Hybrid Inference
par: Sevriugov, Egor, et autres
Publié: (2024)
par: Sevriugov, Egor, et autres
Publié: (2024)
Documents similaires
-
Test-Time Reasoning Through Visual Human Preferences with VLMs and Soft Rewards
par: Gambashidze, Alexander, et autres
Publié: (2025) -
Listener-Rewarded Thinking in VLMs for Image Preferences
par: Gambashidze, Alexander, et autres
Publié: (2025) -
Back to Basics: Revisiting Exploration in Reinforcement Learning for LLM Reasoning via Generative Probabilities
par: Li, Pengyi, et autres
Publié: (2026) -
Confidence Is All You Need: Few-Shot RL Fine-Tuning of Language Models
par: Li, Pengyi, et autres
Publié: (2025) -
OmniFusion Technical Report
par: Goncharova, Elizaveta, et autres
Publié: (2024)