TOGA: Temporally Grounded Open-Ended Video QA with Weak Supervision
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Gupta, Ayush, Roy, Anirban, Chellappa, Rama, Bastian, Nathaniel D., Velasquez, Alvaro, Jha, Susmit |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
MimicGait: A Model Agnostic approach for Occluded Gait Recognition using Correlational Knowledge Distillation
von: Gupta, Ayush, et al.
Veröffentlicht: (2025)
von: Gupta, Ayush, et al.
Veröffentlicht: (2025)
Mind the Gap: Bridging Occlusion in Gait Recognition via Residual Gap Correction
von: Gupta, Ayush, et al.
Veröffentlicht: (2025)
von: Gupta, Ayush, et al.
Veröffentlicht: (2025)
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025)
DiffInf: Influence-Guided Diffusion for Supervision Alignment in Facial Attribute Learning
von: Pal, Basudha, et al.
Veröffentlicht: (2026)
von: Pal, Basudha, et al.
Veröffentlicht: (2026)
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
von: Padhi, Trilok, et al.
Veröffentlicht: (2025)
von: Padhi, Trilok, et al.
Veröffentlicht: (2025)
Zero-Shot Personalization of Objects via Textual Inversion
von: Roy, Aniket, et al.
Veröffentlicht: (2026)
von: Roy, Aniket, et al.
Veröffentlicht: (2026)
GaitContour: Efficient Gait Recognition based on a Contour-Pose Representation
von: Guo, Yuxiang, et al.
Veröffentlicht: (2023)
von: Guo, Yuxiang, et al.
Veröffentlicht: (2023)
MovieRecapsQA: A Multimodal Open-Ended Video Question-Answering Benchmark
von: Shaar, Shaden, et al.
Veröffentlicht: (2026)
von: Shaar, Shaden, et al.
Veröffentlicht: (2026)
Contextual Self-paced Learning for Weakly Supervised Spatio-Temporal Video Grounding
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
von: Kumar, Akash, et al.
Veröffentlicht: (2025)
Template-based Multi-Domain Face Recognition
von: Nanduri, Anirudh, et al.
Veröffentlicht: (2024)
von: Nanduri, Anirudh, et al.
Veröffentlicht: (2024)
CLR-Face: Conditional Latent Refinement for Blind Face Restoration Using Score-Based Diffusion Models
von: Suin, Maitreya, et al.
Veröffentlicht: (2024)
von: Suin, Maitreya, et al.
Veröffentlicht: (2024)
MultLFG: Training-free Multi-LoRA composition using Frequency-domain Guidance
von: Roy, Aniket, et al.
Veröffentlicht: (2025)
von: Roy, Aniket, et al.
Veröffentlicht: (2025)
Pix2Key: Controllable Open-Vocabulary Retrieval with Semantic Decomposition and Self-Supervised Visual Dictionary Learning
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2026)
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2026)
Cross-Modal Fusion and Attention Mechanism for Weakly Supervised Video Anomaly Detection
von: Ghadiya, Ayush, et al.
Veröffentlicht: (2024)
von: Ghadiya, Ayush, et al.
Veröffentlicht: (2024)
ViT-Linearizer: Distilling Quadratic Knowledge into Linear-Time Vision Models
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2025)
von: Wei, Guoyizhe, et al.
Veröffentlicht: (2025)
STPro: Spatial and Temporal Progressive Learning for Weakly Supervised Spatio-Temporal Grounding
von: Garg, Aaryan, et al.
Veröffentlicht: (2025)
von: Garg, Aaryan, et al.
Veröffentlicht: (2025)
Polysemantic Dropout: Conformal OOD Detection for Specialized LLMs
von: Gupta, Ayush, et al.
Veröffentlicht: (2025)
von: Gupta, Ayush, et al.
Veröffentlicht: (2025)
Uncertainty-quantified Pulse Signal Recovery from Facial Video using Regularized Stochastic Interpolants
von: Shenoy, Vineet R., et al.
Veröffentlicht: (2026)
von: Shenoy, Vineet R., et al.
Veröffentlicht: (2026)
Weakly Supervised Temporal Sentence Grounding via Positive Sample Mining
von: Dong, Lu, et al.
Veröffentlicht: (2025)
von: Dong, Lu, et al.
Veröffentlicht: (2025)
A Quantitative Evaluation of the Expressivity of BMI, Pose and Gender in Body Embeddings for Recognition and Identification
von: Pal, Basudha, et al.
Veröffentlicht: (2025)
von: Pal, Basudha, et al.
Veröffentlicht: (2025)
Cross-Spectral Body Recognition with Side Information Embedding: Benchmarks on LLCM and Analyzing Range-Induced Occlusions on IJB-MDF
von: Nanduri, Anirudh, et al.
Veröffentlicht: (2025)
von: Nanduri, Anirudh, et al.
Veröffentlicht: (2025)
Multi-Domain Biometric Recognition using Body Embeddings
von: Nanduri, Anirudh, et al.
Veröffentlicht: (2025)
von: Nanduri, Anirudh, et al.
Veröffentlicht: (2025)
Beyond Isolated Facts: Synthesizing Narrative and Grounded Supervision for VideoQA
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
von: Liang, Jianxin, et al.
Veröffentlicht: (2025)
EtC: Temporal Boundary Expand then Clarify for Weakly Supervised Video Grounding with Multimodal Large Language Model
von: Li, Guozhang, et al.
Veröffentlicht: (2023)
von: Li, Guozhang, et al.
Veröffentlicht: (2023)
VILLS -- Video-Image Learning to Learn Semantics for Person Re-Identification
von: Huang, Siyuan, et al.
Veröffentlicht: (2023)
von: Huang, Siyuan, et al.
Veröffentlicht: (2023)
Video-GroundingDINO: Towards Open-Vocabulary Spatio-Temporal Video Grounding
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
von: Wasim, Syed Talal, et al.
Veröffentlicht: (2023)
NeuS-QA: Grounding Long-Form Video Understanding in Temporal Logic and Neuro-Symbolic Reasoning
von: Shah, Sahil, et al.
Veröffentlicht: (2025)
von: Shah, Sahil, et al.
Veröffentlicht: (2025)
StimuVAR: Spatiotemporal Stimuli-aware Video Affective Reasoning with Multimodal Large Language Models
von: Guo, Yuxiang, et al.
Veröffentlicht: (2024)
von: Guo, Yuxiang, et al.
Veröffentlicht: (2024)
Siamese Learning with Joint Alignment and Regression for Weakly-Supervised Video Paragraph Grounding
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
von: Tan, Chaolei, et al.
Veröffentlicht: (2024)
Video-in-the-Loop: Span-Grounded Long Video QA with Interleaved Reasoning
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
von: Wang, Chendong, et al.
Veröffentlicht: (2025)
Grounded Multi-Hop VideoQA in Long-Form Egocentric Videos
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
von: Chen, Qirui, et al.
Veröffentlicht: (2024)
Finding Optimal Video Moment without Training: Gaussian Boundary Optimization for Weakly Supervised Video Grounding
von: Kim, Sunoh, et al.
Veröffentlicht: (2026)
von: Kim, Sunoh, et al.
Veröffentlicht: (2026)
SPIDER: Spatial Image CorresponDence Estimator for Robust Calibration
von: Shao, Zhimin, et al.
Veröffentlicht: (2025)
von: Shao, Zhimin, et al.
Veröffentlicht: (2025)
CAM3R: Camera-Agnostic Model for 3D Reconstruction
von: Guruprasad, Namitha, et al.
Veröffentlicht: (2026)
von: Guruprasad, Namitha, et al.
Veröffentlicht: (2026)
Language-guided Open-world Video Anomaly Detection under Weak Supervision
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
von: Liu, Zihao, et al.
Veröffentlicht: (2025)
RANKVIDEO: Reasoning Reranking for Text-to-Video Retrieval
von: Skow, Tyler, et al.
Veröffentlicht: (2026)
von: Skow, Tyler, et al.
Veröffentlicht: (2026)
Ranking Distillation for Open-Ended Video Question Answering with Insufficient Labels
von: Liang, Tianming, et al.
Veröffentlicht: (2024)
von: Liang, Tianming, et al.
Veröffentlicht: (2024)
Gaussian Mixture Proposals with Pull-Push Learning Scheme to Capture Diverse Events for Weakly Supervised Temporal Video Grounding
von: Kim, Sunoh, et al.
Veröffentlicht: (2023)
von: Kim, Sunoh, et al.
Veröffentlicht: (2023)
TriFusion-AE: Language-Guided Depth and LiDAR Fusion for Robust Point Cloud Processing
von: Neogi, Susmit
Veröffentlicht: (2025)
von: Neogi, Susmit
Veröffentlicht: (2025)
MV2MAE: Multi-View Video Masked Autoencoders
von: Shah, Ketul, et al.
Veröffentlicht: (2024)
von: Shah, Ketul, et al.
Veröffentlicht: (2024)
Ähnliche Einträge
-
MimicGait: A Model Agnostic approach for Occluded Gait Recognition using Correlational Knowledge Distillation
von: Gupta, Ayush, et al.
Veröffentlicht: (2025) -
Mind the Gap: Bridging Occlusion in Gait Recognition via Residual Gap Correction
von: Gupta, Ayush, et al.
Veröffentlicht: (2025) -
Enrich and Detect: Video Temporal Grounding with Multimodal LLMs
von: Pramanick, Shraman, et al.
Veröffentlicht: (2025) -
DiffInf: Influence-Guided Diffusion for Supervision Alignment in Facial Attribute Learning
von: Pal, Basudha, et al.
Veröffentlicht: (2026) -
Calibrating Uncertainty Quantification of Multi-Modal LLMs using Grounding
von: Padhi, Trilok, et al.
Veröffentlicht: (2025)