Overcoming Semantic Dilution in Transformer-Based Next Frame Prediction
Fuente:
arXiv
Guardado en:
| Autores principales: | Nguyen, Hy, Thudumu, Srikanth, Du, Hung, Vasa, Rajesh, Mouzakis, Kon |
|---|---|
| Formato: | Preprint |
| Publicado: |
2025
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
CSAOT: Cooperative Multi-Agent System for Active Object Tracking
por: Nguyen, Hy, et al.
Publicado: (2025)
por: Nguyen, Hy, et al.
Publicado: (2025)
Contextual Knowledge Sharing in Multi-Agent Reinforcement Learning with Decentralized Communication and Coordination
por: Du, Hung, et al.
Publicado: (2025)
por: Du, Hung, et al.
Publicado: (2025)
Goal-Oriented Multi-Agent Reinforcement Learning for Decentralized Agent Teams
por: Du, Hung, et al.
Publicado: (2025)
por: Du, Hung, et al.
Publicado: (2025)
A Survey on Context-Aware Multi-Agent Systems: Techniques, Challenges and Future Directions
por: Du, Hung, et al.
Publicado: (2024)
por: Du, Hung, et al.
Publicado: (2024)
Local Control Networks (LCNs): Optimizing Flexibility in Neural Network Data Pattern Capture
por: Nguyen, Hy, et al.
Publicado: (2025)
por: Nguyen, Hy, et al.
Publicado: (2025)
Dual-Branch HNSW Approach with Skip Bridges and LID-Driven Optimization
por: Nguyen, Hy, et al.
Publicado: (2025)
por: Nguyen, Hy, et al.
Publicado: (2025)
The M-factor: A Novel Metric for Evaluating Neural Architecture Search in Resource-Constrained Environments
por: Thudumu, Srikanth, et al.
Publicado: (2025)
por: Thudumu, Srikanth, et al.
Publicado: (2025)
Pretraining Objective Matters in Extreme Low-Data FGVC: A Backbone-Controlled Study
por: Hackett, Alexander, et al.
Publicado: (2026)
por: Hackett, Alexander, et al.
Publicado: (2026)
Playing with Transformer at 30+ FPS via Next-Frame Diffusion
por: Cheng, Xinle, et al.
Publicado: (2025)
por: Cheng, Xinle, et al.
Publicado: (2025)
VideoAR: Autoregressive Video Generation via Next-Frame & Scale Prediction
por: Ji, Longbin, et al.
Publicado: (2026)
por: Ji, Longbin, et al.
Publicado: (2026)
XAI-Enhanced Semantic Segmentation Models for Visual Quality Inspection
por: Clement, Tobias, et al.
Publicado: (2024)
por: Clement, Tobias, et al.
Publicado: (2024)
Explanation-Driven Counterfactual Testing for Faithfulness in Vision-Language Model Explanations
por: Ding, Sihao, et al.
Publicado: (2025)
por: Ding, Sihao, et al.
Publicado: (2025)
Novel 3D Binary Indexed Tree for Volume Computation of 3D Reconstructed Models from Volumetric Data
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024)
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024)
Overcoming the Curvature Bottleneck in MeanFlow
por: Zhang, Xinxi, et al.
Publicado: (2025)
por: Zhang, Xinxi, et al.
Publicado: (2025)
Fostering Video Reasoning via Next-Event Prediction
por: Wang, Haonan, et al.
Publicado: (2025)
por: Wang, Haonan, et al.
Publicado: (2025)
RotCAtt-TransUNet++: Novel Deep Neural Network for Sophisticated Cardiac Segmentation
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024)
por: Nguyen-Le, Quoc-Bao, et al.
Publicado: (2024)
Predicting the Next Action by Modeling the Abstract Goal
por: Roy, Debaditya, et al.
Publicado: (2022)
por: Roy, Debaditya, et al.
Publicado: (2022)
SwiftTry: Fast and Consistent Video Virtual Try-On with Diffusion Models
por: Nguyen, Hung, et al.
Publicado: (2024)
por: Nguyen, Hung, et al.
Publicado: (2024)
Enhancing the Fairness and Performance of Edge Cameras with Explainable AI
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
por: Nguyen, Truong Thanh Hung, et al.
Publicado: (2024)
Next Best View Selections for Semantic and Dynamic 3D Gaussian Splatting
por: Li, Yiqian, et al.
Publicado: (2025)
por: Li, Yiqian, et al.
Publicado: (2025)
Semantic Causality-Aware Vision-Based 3D Occupancy Prediction
por: Chen, Dubing, et al.
Publicado: (2025)
por: Chen, Dubing, et al.
Publicado: (2025)
SDGOCC: Semantic and Depth-Guided Bird's-Eye View Transformation for 3D Multimodal Occupancy Prediction
por: Duan, Zaipeng, et al.
Publicado: (2025)
por: Duan, Zaipeng, et al.
Publicado: (2025)
GPT-4V Takes the Wheel: Promises and Challenges for Pedestrian Behavior Prediction
por: Huang, Jia, et al.
Publicado: (2023)
por: Huang, Jia, et al.
Publicado: (2023)
Next Block Prediction: Video Generation via Semi-Autoregressive Modeling
por: Ren, Shuhuai, et al.
Publicado: (2025)
por: Ren, Shuhuai, et al.
Publicado: (2025)
Efficient and Concise Explanations for Object Detection with Gaussian-Class Activation Mapping Explainer
por: Nguyen, Quoc Khanh, et al.
Publicado: (2024)
por: Nguyen, Quoc Khanh, et al.
Publicado: (2024)
Power of Boundary and Reflection: Semantic Transparent Object Segmentation using Pyramid Vision Transformer with Transparent Cues
por: Vu, Tuan-Anh, et al.
Publicado: (2025)
por: Vu, Tuan-Anh, et al.
Publicado: (2025)
GAIS: Frame-Level Gated Audio-Visual Integration with Semantic Variance-Scaled Perturbation for Text-Video Retrieval
por: Yang, Bowen, et al.
Publicado: (2025)
por: Yang, Bowen, et al.
Publicado: (2025)
GaussianFormer: Scene as Gaussians for Vision-Based 3D Semantic Occupancy Prediction
por: Huang, Yuanhui, et al.
Publicado: (2024)
por: Huang, Yuanhui, et al.
Publicado: (2024)
Representation Separation for Semantic Segmentation with Vision Transformers
por: Hong, Yuanduo, et al.
Publicado: (2022)
por: Hong, Yuanduo, et al.
Publicado: (2022)
DisBeaNet: A Deep Neural Network to augment Unmanned Surface Vessels for maritime situational awareness
por: Vemula, Srikanth, et al.
Publicado: (2024)
por: Vemula, Srikanth, et al.
Publicado: (2024)
Transfusion: Predict the Next Token and Diffuse Images with One Multi-Modal Model
por: Zhou, Chunting, et al.
Publicado: (2024)
por: Zhou, Chunting, et al.
Publicado: (2024)
Visual Autoregressive Modeling: Scalable Image Generation via Next-Scale Prediction
por: Tian, Keyu, et al.
Publicado: (2024)
por: Tian, Keyu, et al.
Publicado: (2024)
Improving Zero-Shot Object-Level Change Detection by Incorporating Visual Correspondence
por: Nguyen, Hung Huy, et al.
Publicado: (2025)
por: Nguyen, Hung Huy, et al.
Publicado: (2025)
M-LLM Based Video Frame Selection for Efficient Video Understanding
por: Hu, Kai, et al.
Publicado: (2025)
por: Hu, Kai, et al.
Publicado: (2025)
FrameMind: Frame-Interleaved Video Reasoning via Reinforcement Learning
por: Ge, Haonan, et al.
Publicado: (2025)
por: Ge, Haonan, et al.
Publicado: (2025)
GraphGSOcc: Semantic-Geometric Graph Transformer with Dynamic-Static Decoupling for 3D Gaussian Splatting-based Occupancy Prediction
por: Song, Ke, et al.
Publicado: (2025)
por: Song, Ke, et al.
Publicado: (2025)
DiffAttn: Diffusion-Based Drivers' Visual Attention Prediction with LLM-Enhanced Semantic Reasoning
por: Liu, Weimin, et al.
Publicado: (2026)
por: Liu, Weimin, et al.
Publicado: (2026)
SOccDPT: Semi-Supervised 3D Semantic Occupancy from Dense Prediction Transformers trained under memory constraints
por: Ganesh, Aditya Nalgunda
Publicado: (2023)
por: Ganesh, Aditya Nalgunda
Publicado: (2023)
KEPT: Knowledge-Enhanced Prediction of Trajectories from Consecutive Driving Frames with Vision-Language Models
por: Wang, Yujin, et al.
Publicado: (2025)
por: Wang, Yujin, et al.
Publicado: (2025)
ECMNet:Lightweight Semantic Segmentation with Efficient CNN-Mamba Network
por: Du, Feixiang, et al.
Publicado: (2025)
por: Du, Feixiang, et al.
Publicado: (2025)
Ejemplares similares
-
CSAOT: Cooperative Multi-Agent System for Active Object Tracking
por: Nguyen, Hy, et al.
Publicado: (2025) -
Contextual Knowledge Sharing in Multi-Agent Reinforcement Learning with Decentralized Communication and Coordination
por: Du, Hung, et al.
Publicado: (2025) -
Goal-Oriented Multi-Agent Reinforcement Learning for Decentralized Agent Teams
por: Du, Hung, et al.
Publicado: (2025) -
A Survey on Context-Aware Multi-Agent Systems: Techniques, Challenges and Future Directions
por: Du, Hung, et al.
Publicado: (2024) -
Local Control Networks (LCNs): Optimizing Flexibility in Neural Network Data Pattern Capture
por: Nguyen, Hy, et al.
Publicado: (2025)