SGCap: Decoding Semantic Group for Zero-shot Video Captioning
Fuente:
arXiv
Saved in:
| Main Authors: | Pan, Zeyu, Li, Ping, Wang, Wenxiao |
|---|---|
| Format: | Preprint |
| Published: |
2025
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
by: Han, Zeyu, et al.
Published: (2023)
by: Han, Zeyu, et al.
Published: (2023)
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024)
by: Zeng, Zequn, et al.
Published: (2024)
Few-shot Semantic Encoding and Decoding for Video Surveillance
by: Cheng, Baoping, et al.
Published: (2025)
by: Cheng, Baoping, et al.
Published: (2025)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
by: Lu, Yifan, et al.
Published: (2023)
by: Lu, Yifan, et al.
Published: (2023)
Zero-shot Reconstruction of In-Scene Object Manipulation from Video
by: Lin, Dixuan, et al.
Published: (2025)
by: Lin, Dixuan, et al.
Published: (2025)
Split Matching for Inductive Zero-shot Semantic Segmentation
by: Chen, Jialei, et al.
Published: (2025)
by: Chen, Jialei, et al.
Published: (2025)
Pseudo-labeling with Keyword Refining for Few-Supervised Video Captioning
by: Li, Ping, et al.
Published: (2024)
by: Li, Ping, et al.
Published: (2024)
Knowledge Guided Entity-aware Video Captioning and A Basketball Benchmark
by: Xi, Zeyu, et al.
Published: (2024)
by: Xi, Zeyu, et al.
Published: (2024)
Dense Video Captioning Using Unsupervised Semantic Information
by: Estevam, Valter, et al.
Published: (2021)
by: Estevam, Valter, et al.
Published: (2021)
Unbiased Semantic Decoding with Vision Foundation Models for Few-shot Segmentation
by: Wang, Jin, et al.
Published: (2025)
by: Wang, Jin, et al.
Published: (2025)
Sample-level Adaptive Knowledge Distillation for Action Recognition
by: Li, Ping, et al.
Published: (2025)
by: Li, Ping, et al.
Published: (2025)
ViDiC: Video Difference Captioning
by: Wu, Jiangtao, et al.
Published: (2025)
by: Wu, Jiangtao, et al.
Published: (2025)
Attention Based Encoder Decoder Model for Video Captioning in Nepali (2023)
by: Parajuli, Kabita, et al.
Published: (2023)
by: Parajuli, Kabita, et al.
Published: (2023)
Prioritized Semantic Learning for Zero-shot Instance Navigation
by: Sun, Xinyu, et al.
Published: (2024)
by: Sun, Xinyu, et al.
Published: (2024)
IG Captioner: Information Gain Captioners are Strong Zero-shot Classifiers
by: Yang, Chenglin, et al.
Published: (2023)
by: Yang, Chenglin, et al.
Published: (2023)
RETTA: Retrieval-Enhanced Test-Time Adaptation for Zero-Shot Video Captioning
by: Ma, Yunchuan, et al.
Published: (2024)
by: Ma, Yunchuan, et al.
Published: (2024)
AlignZeg: Mitigating Objective Misalignment for Zero-shot Semantic Segmentation
by: Ge, Jiannan, et al.
Published: (2024)
by: Ge, Jiannan, et al.
Published: (2024)
Generalizable Semantic Vision Query Generation for Zero-shot Panoptic and Semantic Segmentation
by: Chen, Jialei, et al.
Published: (2024)
by: Chen, Jialei, et al.
Published: (2024)
VideoMaker: Zero-shot Customized Video Generation with the Inherent Force of Video Diffusion Models
by: Wu, Tao, et al.
Published: (2024)
by: Wu, Tao, et al.
Published: (2024)
Attribute Distribution Modeling and Semantic-Visual Alignment for Generative Zero-shot Learning
by: Pu, Haojie, et al.
Published: (2026)
by: Pu, Haojie, et al.
Published: (2026)
Combating Semantic Contamination in Learning with Label Noise
by: Fan, Wenxiao, et al.
Published: (2024)
by: Fan, Wenxiao, et al.
Published: (2024)
PSVMA+: Exploring Multi-granularity Semantic-visual Adaption for Generalized Zero-shot Learning
by: Liu, Man, et al.
Published: (2024)
by: Liu, Man, et al.
Published: (2024)
The Devil is in the Distributions: Explicit Modeling of Scene Content is Key in Zero-Shot Video Captioning
by: Tian, Mingkai, et al.
Published: (2025)
by: Tian, Mingkai, et al.
Published: (2025)
TPCap: Unlocking Zero-Shot Image Captioning with Trigger-Augmented and Multi-Modal Purification Modules
by: Zhang, Ruoyu, et al.
Published: (2025)
by: Zhang, Ruoyu, et al.
Published: (2025)
Bridge the Gap Between Visual and Linguistic Comprehension for Generalized Zero-shot Semantic Segmentation
by: Guo, Xiaoqing, et al.
Published: (2025)
by: Guo, Xiaoqing, et al.
Published: (2025)
Bridged Semantic Alignment for Zero-shot 3D Medical Image Diagnosis
by: Lai, Haoran, et al.
Published: (2025)
by: Lai, Haoran, et al.
Published: (2025)
Encoder-Decoder Based Long Short-Term Memory (LSTM) Model for Video Captioning
by: Adewale, Sikiru, et al.
Published: (2023)
by: Adewale, Sikiru, et al.
Published: (2023)
Partial CLIP is Enough: Chimera-Seg for Zero-shot Semantic Segmentation
by: Chen, Jialei, et al.
Published: (2025)
by: Chen, Jialei, et al.
Published: (2025)
CLIPS: An Enhanced CLIP Framework for Learning with Synthetic Captions
by: Liu, Yanqing, et al.
Published: (2024)
by: Liu, Yanqing, et al.
Published: (2024)
Captioning for Text-Video Retrieval via Dual-Group Direct Preference Optimization
by: Lee, Ji Soo, et al.
Published: (2025)
by: Lee, Ji Soo, et al.
Published: (2025)
Depth-aware Test-Time Training for Zero-shot Video Object Segmentation
by: Liu, Weihuang, et al.
Published: (2024)
by: Liu, Weihuang, et al.
Published: (2024)
ShareGPT4Video: Improving Video Understanding and Generation with Better Captions
by: Chen, Lin, et al.
Published: (2024)
by: Chen, Lin, et al.
Published: (2024)
Exploiting Auxiliary Caption for Video Grounding
by: Li, Hongxiang, et al.
Published: (2023)
by: Li, Hongxiang, et al.
Published: (2023)
VideoINSTA: Zero-shot Long Video Understanding via Informative Spatial-Temporal Reasoning with LLMs
by: Liao, Ruotong, et al.
Published: (2024)
by: Liao, Ruotong, et al.
Published: (2024)
Live Video Captioning
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
by: Blanco-Fernández, Eduardo, et al.
Published: (2024)
Knowledge NeRF: Few-shot Novel View Synthesis for Dynamic Articulated Objects
by: Cai, Wenxiao, et al.
Published: (2024)
by: Cai, Wenxiao, et al.
Published: (2024)
EVA: Zero-shot Accurate Attributes and Multi-Object Video Editing
by: Yang, Xiangpeng, et al.
Published: (2024)
by: Yang, Xiangpeng, et al.
Published: (2024)
Adaptive Multi-source Predictor for Zero-shot Video Object Segmentation
by: Zhao, Xiaoqi, et al.
Published: (2023)
by: Zhao, Xiaoqi, et al.
Published: (2023)
Zero-shot Prompt-based Video Encoder for Surgical Gesture Recognition
by: Rao, Mingxing, et al.
Published: (2024)
by: Rao, Mingxing, et al.
Published: (2024)
Semantic Segmentation of Transparent and Opaque Drinking Glasses with the Help of Zero-shot Learning
by: Blänsdorf, Annalena, et al.
Published: (2025)
by: Blänsdorf, Annalena, et al.
Published: (2025)
Similar Items
-
Zero-shot Referring Expression Comprehension via Structural Similarity Between Images and Captions
by: Han, Zeyu, et al.
Published: (2023) -
MeaCap: Memory-Augmented Zero-shot Image Captioning
by: Zeng, Zequn, et al.
Published: (2024) -
Few-shot Semantic Encoding and Decoding for Video Surveillance
by: Cheng, Baoping, et al.
Published: (2025) -
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
by: Lu, Yifan, et al.
Published: (2023) -
Zero-shot Reconstruction of In-Scene Object Manipulation from Video
by: Lin, Dixuan, et al.
Published: (2025)