Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
Fuente:
arXiv
Guardado en:
| Autores principales: | Anh, Duy Le Dinh, Tran, Kim Hoang, Nguyen, Quang-Thuc, Le, Ngan Hoang |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
por: Li, Danyang, et al.
Publicado: (2025)
por: Li, Danyang, et al.
Publicado: (2025)
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
por: Le, Van-Truong
Publicado: (2026)
por: Le, Van-Truong
Publicado: (2026)
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
por: Lim, Shoon Kit, et al.
Publicado: (2025)
por: Lim, Shoon Kit, et al.
Publicado: (2025)
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)
por: Dai, Song, et al.
Publicado: (2025)
Universal Adversarial Attack on Aligned Multimodal LLMs
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
por: Rahmatullaev, Temurbek, et al.
Publicado: (2025)
Game-RL: Synthesizing Multimodal Verifiable Game Data to Boost VLMs' General Reasoning
por: Tong, Jingqi, et al.
Publicado: (2025)
por: Tong, Jingqi, et al.
Publicado: (2025)
Memory-Efficient Differentially Private Training with Gradient Random Projection
por: Mulrooney, Alex, et al.
Publicado: (2025)
por: Mulrooney, Alex, et al.
Publicado: (2025)
PhysNote: Self-Knowledge Notes for Evolvable Physical Reasoning in Vision-Language Model
por: Zhang, Sinin, et al.
Publicado: (2026)
por: Zhang, Sinin, et al.
Publicado: (2026)
ICG: Improving Cover Image Generation via MLLM-based Prompting and Personalized Preference Alignment
por: Bian, Zhipeng, et al.
Publicado: (2026)
por: Bian, Zhipeng, et al.
Publicado: (2026)
Learning the meanings of function words from grounded language using a visual question answering model
por: Portelance, Eva, et al.
Publicado: (2023)
por: Portelance, Eva, et al.
Publicado: (2023)
Physics-R1: An Audited Olympiad Corpus and Recipe for Visual Physics Reasoning
por: Yang, Shan
Publicado: (2026)
por: Yang, Shan
Publicado: (2026)
Spatially-Aware Speaker for Vision-and-Language Navigation Instruction Generation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024)
Evaluating Voice Command Pipelines for Drone Control: From STT and LLM to Direct Classification and Siamese Networks
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
por: Simões, Lucca Emmanuel Pineli, et al.
Publicado: (2024)
GroundCap: A Visually Grounded Image Captioning Dataset
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Correspondence of high-dimensional emotion structures elicited by video clips between humans and Multimodal LLMs
por: Asanuma, Haruka, et al.
Publicado: (2025)
por: Asanuma, Haruka, et al.
Publicado: (2025)
HuMoCon: Concept Discovery for Human Motion Understanding
por: Fang, Qihang, et al.
Publicado: (2025)
por: Fang, Qihang, et al.
Publicado: (2025)
Survey Transfer Learning: Recycling Data with Silicon Responses
por: Amini, Ali
Publicado: (2025)
por: Amini, Ali
Publicado: (2025)
DriveMRP: Enhancing Vision-Language Models with Synthetic Motion Data for Motion Risk Prediction
por: Hou, Zhiyi, et al.
Publicado: (2025)
por: Hou, Zhiyi, et al.
Publicado: (2025)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
por: Shah, Nisarg A., et al.
Publicado: (2025)
por: Shah, Nisarg A., et al.
Publicado: (2025)
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
por: Bucher, Martin JJ., et al.
Publicado: (2025)
por: Bucher, Martin JJ., et al.
Publicado: (2025)
A Closer Look at Bias and Chain-of-Thought Faithfulness of Large (Vision) Language Models
por: Balasubramanian, Sriram, et al.
Publicado: (2025)
por: Balasubramanian, Sriram, et al.
Publicado: (2025)
MIRA: Empowering One-Touch AI Services on Smartphones with MLLM-based Instruction Recommendation
por: Bian, Zhipeng, et al.
Publicado: (2025)
por: Bian, Zhipeng, et al.
Publicado: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
por: Cui, Shaoyang, et al.
Publicado: (2026)
por: Cui, Shaoyang, et al.
Publicado: (2026)
Defending against Backdoor Attacks via Module Switching
por: Li, Weijun, et al.
Publicado: (2025)
por: Li, Weijun, et al.
Publicado: (2025)
What Makes a Good Story and How Can We Measure It? A Comprehensive Survey of Story Evaluation
por: Yang, Dingyi, et al.
Publicado: (2024)
por: Yang, Dingyi, et al.
Publicado: (2024)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
por: Oliveira, Daniel A. P., et al.
Publicado: (2025)
Leum-VL Technical Report
por: He, Yuxuan, et al.
Publicado: (2026)
por: He, Yuxuan, et al.
Publicado: (2026)
Reframing linguistic bootstrapping as joint inference using visually-grounded grammar induction models
por: Portelance, Eva, et al.
Publicado: (2024)
por: Portelance, Eva, et al.
Publicado: (2024)
Enhancing Sports Strategy with Video Analytics and Data Mining: Assessing the effectiveness of Multimodal LLMs in tennis video analysis
por: Teo, Charlton
Publicado: (2025)
por: Teo, Charlton
Publicado: (2025)
Human-Robot Dialogue Annotation for Multi-Modal Common Ground
por: Bonial, Claire, et al.
Publicado: (2024)
por: Bonial, Claire, et al.
Publicado: (2024)
Beyond Localization: A Comprehensive Diagnosis of Perspective-Conditioned Spatial Reasoning in MLLMs from Omnidirectional Images
por: Chen, Yuangong, et al.
Publicado: (2026)
por: Chen, Yuangong, et al.
Publicado: (2026)
TensLoRA: Tensor Alternatives for Low-Rank Adaptation
por: Marmoret, Axel, et al.
Publicado: (2025)
por: Marmoret, Axel, et al.
Publicado: (2025)
WildRoadBench: A Wild Aerial Road-Damage Grounding Benchmark for Vision-Language Models and Autonomous Agents
por: Liu, Bingnan, et al.
Publicado: (2026)
por: Liu, Bingnan, et al.
Publicado: (2026)
Anonymization-Enhanced Privacy Protection for Mobile GUI Agents: Available but Invisible
por: Zhao, Lepeng, et al.
Publicado: (2026)
por: Zhao, Lepeng, et al.
Publicado: (2026)
Waking Up Blind: Cold-Start Optimization of Supervision-Free Agentic Trajectories for Grounded Visual Perception
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
por: Bajpai, Ashutosh, et al.
Publicado: (2026)
Unpacking Failure Modes of Generative Policies: Runtime Monitoring of Consistency and Progress
por: Agia, Christopher, et al.
Publicado: (2024)
por: Agia, Christopher, et al.
Publicado: (2024)
EMOS: Embodiment-aware Heterogeneous Multi-robot Operating System with LLM Agents
por: Chen, Junting, et al.
Publicado: (2024)
por: Chen, Junting, et al.
Publicado: (2024)
MM-Conv: A Multimodal Dataset and Benchmark for Context-Aware Grounding in 3D Dialogue
por: Deichler, Anna, et al.
Publicado: (2026)
por: Deichler, Anna, et al.
Publicado: (2026)
K-MetBench: A Multi-Dimensional Benchmark for Fine-Grained Evaluation of Expert Reasoning, Locality, and Multimodality in Meteorology
por: Kim, Soyeon, et al.
Publicado: (2026)
por: Kim, Soyeon, et al.
Publicado: (2026)
Ejemplares similares
-
OpenMap: Instruction Grounding via Open-Vocabulary Visual-Language Mapping
por: Li, Danyang, et al.
Publicado: (2025) -
From Benchmarking to Reasoning: A Dual-Aspect, Large-Scale Evaluation of LLMs on Vietnamese Legal Text
por: Le, Van-Truong
Publicado: (2026) -
Taking Flight with Dialogue: Enabling Natural Language Control for PX4-based Drone Agent
por: Lim, Shoon Kit, et al.
Publicado: (2025) -
StratXplore: Strategic Novelty-seeking and Instruction-aligned Exploration for Vision and Language Navigation
por: Gopinathan, Muraleekrishna, et al.
Publicado: (2024) -
PhysicsArena: The First Multimodal Physics Reasoning Benchmark Exploring Variable, Process, and Solution Dimensions
por: Dai, Song, et al.
Publicado: (2025)