Guardado en:
| Autores principales: | Sugihara, Tomoya, Masuda, Shuntaro, Xiao, Ling, Yamasaki, Toshihiko |
|---|---|
| Formato: | Preprint |
| Publicado: |
2024
|
| Materias: | |
| Acceso en línea: | https://arxiv.org/abs/2405.08890 |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
por: Xiao, Ling, et al.
Publicado: (2022)
por: Xiao, Ling, et al.
Publicado: (2022)
A Multihead Continual Learning Framework for Fine-Grained Fashion Image Retrieval with Contrastive Learning and Exponential Moving Average Distillation
por: Xiao, Ling, et al.
Publicado: (2026)
por: Xiao, Ling, et al.
Publicado: (2026)
Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation
por: Tanji, Naoto, et al.
Publicado: (2025)
por: Tanji, Naoto, et al.
Publicado: (2025)
Language-Guided Graph Representation Learning for Video Summarization
por: Li, Wenrui, et al.
Publicado: (2025)
por: Li, Wenrui, et al.
Publicado: (2025)
Language-guided Detection and Mitigation of Unknown Dataset Bias
por: Zhao, Zaiying, et al.
Publicado: (2024)
por: Zhao, Zaiying, et al.
Publicado: (2024)
Online Open-set Semi-supervised Object Detection with Dual Competing Head
por: Wang, Zerun, et al.
Publicado: (2023)
por: Wang, Zerun, et al.
Publicado: (2023)
Spectral Probing of Feature Upsamplers in 2D-to-3D Scene Reconstruction
por: Xiao, Ling, et al.
Publicado: (2026)
por: Xiao, Ling, et al.
Publicado: (2026)
Face2Diffusion for Fast and Editable Face Personalization
por: Shiohara, Kaede, et al.
Publicado: (2024)
por: Shiohara, Kaede, et al.
Publicado: (2024)
Unified Vector Floorplan Generation via Markup Representation
por: Shiohara, Kaede, et al.
Publicado: (2026)
por: Shiohara, Kaede, et al.
Publicado: (2026)
CE-FAM: Concept-Based Explanation via Fusion of Activation Maps
por: Kuroki, Michihiro, et al.
Publicado: (2025)
por: Kuroki, Michihiro, et al.
Publicado: (2025)
BSED: Baseline Shapley-Based Explainable Detector
por: Kuroki, Michihiro, et al.
Publicado: (2023)
por: Kuroki, Michihiro, et al.
Publicado: (2023)
Bias Beyond Demographics: Probing Decision Boundaries in Black-Box LVLMs via Counterfactual VQA
por: Zhao, Zaiying, et al.
Publicado: (2025)
por: Zhao, Zaiying, et al.
Publicado: (2025)
Prompts to Summaries: Zero-Shot Language-Guided Video Summarization with Large Language and Video Models
por: Barbara, Mario, et al.
Publicado: (2025)
por: Barbara, Mario, et al.
Publicado: (2025)
Set Prediction Guided by Semantic Concepts for Diverse Video Captioning
por: Lu, Yifan, et al.
Publicado: (2023)
por: Lu, Yifan, et al.
Publicado: (2023)
Realizing Video Summarization from the Path of Language-based Semantic Understanding
por: Mu, Kuan-Chen, et al.
Publicado: (2024)
por: Mu, Kuan-Chen, et al.
Publicado: (2024)
Reward Incremental Learning in Text-to-Image Generation
por: Wang, Maorong, et al.
Publicado: (2024)
por: Wang, Maorong, et al.
Publicado: (2024)
TRIM: A Self-Supervised Video Summarization Framework Maximizing Temporal Relative Information and Representativeness
por: Mishra, Pritam, et al.
Publicado: (2025)
por: Mishra, Pritam, et al.
Publicado: (2025)
SCOMatch: Alleviating Overtrusting in Open-set Semi-supervised Learning
por: Wang, Zerun, et al.
Publicado: (2024)
por: Wang, Zerun, et al.
Publicado: (2024)
Video Summarization with Large Language Models
por: Lee, Min Jung, et al.
Publicado: (2025)
por: Lee, Min Jung, et al.
Publicado: (2025)
Auto-Comp: An Automated Pipeline for Scalable Compositional Probing of Contrastive Vision-Language Models
por: Sbrolli, Cristian, et al.
Publicado: (2026)
por: Sbrolli, Cristian, et al.
Publicado: (2026)
Scene Summarization: Clustering Scene Videos into Spatially Diverse Frames
por: Chen, Chao, et al.
Publicado: (2023)
por: Chen, Chao, et al.
Publicado: (2023)
ExposeAnyone: Personalized Audio-to-Expression Diffusion Models Are Robust Zero-Shot Face Forgery Detectors
por: Shiohara, Kaede, et al.
Publicado: (2026)
por: Shiohara, Kaede, et al.
Publicado: (2026)
ControlVP: Interactive Geometric Refinement of AI-Generated Images with Consistent Vanishing Points
por: Okumura, Ryota, et al.
Publicado: (2025)
por: Okumura, Ryota, et al.
Publicado: (2025)
VideoSAVi: Self-Aligned Video Language Models without Human Supervision
por: Kulkarni, Yogesh, et al.
Publicado: (2024)
por: Kulkarni, Yogesh, et al.
Publicado: (2024)
TRIMMER: A New Paradigm for Video Summarization through Self-Supervised Reinforcement Learning
por: Mishra, Pritam, et al.
Publicado: (2026)
por: Mishra, Pritam, et al.
Publicado: (2026)
VideoSSR: Video Self-Supervised Reinforcement Learning
por: He, Zefeng, et al.
Publicado: (2025)
por: He, Zefeng, et al.
Publicado: (2025)
FILS: Self-Supervised Video Feature Prediction In Semantic Language Space
por: Ahmadian, Mona, et al.
Publicado: (2024)
por: Ahmadian, Mona, et al.
Publicado: (2024)
Rewrite Caption Semantics: Bridging Semantic Gaps for Language-Supervised Semantic Segmentation
por: Xing, Yun, et al.
Publicado: (2023)
por: Xing, Yun, et al.
Publicado: (2023)
RefineVAD: Semantic-Guided Feature Recalibration for Weakly Supervised Video Anomaly Detection
por: Lee, Junhee, et al.
Publicado: (2025)
por: Lee, Junhee, et al.
Publicado: (2025)
InternVideo-Next: Towards General Video Foundation Models without Video-Text Supervision
por: Wang, Chenting, et al.
Publicado: (2025)
por: Wang, Chenting, et al.
Publicado: (2025)
Enhancing Video Memorability Prediction with Text-Motion Cross-modal Contrastive Loss and Its Application in Video Summarization
por: Zhu, Zhiyi, et al.
Publicado: (2025)
por: Zhu, Zhiyi, et al.
Publicado: (2025)
Theoretical Understanding of Learning from Adversarial Perturbations
por: Kumano, Soichiro, et al.
Publicado: (2024)
por: Kumano, Soichiro, et al.
Publicado: (2024)
Wide Two-Layer Networks can Learn from Adversarial Perturbations
por: Kumano, Soichiro, et al.
Publicado: (2024)
por: Kumano, Soichiro, et al.
Publicado: (2024)
Text-Guided Video Masked Autoencoder
por: Fan, David, et al.
Publicado: (2024)
por: Fan, David, et al.
Publicado: (2024)
Adversarial Training from Mean Field Perspective
por: Kumano, Soichiro, et al.
Publicado: (2025)
por: Kumano, Soichiro, et al.
Publicado: (2025)
Adversarially Pretrained Transformers May Be Universally Robust In-Context Learners
por: Kumano, Soichiro, et al.
Publicado: (2025)
por: Kumano, Soichiro, et al.
Publicado: (2025)
VideoClusterNet: Self-Supervised and Adaptive Face Clustering For Videos
por: Walawalkar, Devesh, et al.
Publicado: (2024)
por: Walawalkar, Devesh, et al.
Publicado: (2024)
SelfHVD: Self-Supervised Handheld Video Deblurring
por: Xu, Honglei, et al.
Publicado: (2025)
por: Xu, Honglei, et al.
Publicado: (2025)
Scaling Up Video Summarization Pretraining with Large Language Models
por: Argaw, Dawit Mureja, et al.
Publicado: (2024)
por: Argaw, Dawit Mureja, et al.
Publicado: (2024)
Difficulty Controlled Diffusion Model for Synthesizing Effective Training Data
por: Wang, Zerun, et al.
Publicado: (2024)
por: Wang, Zerun, et al.
Publicado: (2024)
Ejemplares similares
-
Attribute-Guided Multi-Level Attention Network for Fine-Grained Fashion Retrieval
por: Xiao, Ling, et al.
Publicado: (2022) -
A Multihead Continual Learning Framework for Fine-Grained Fashion Image Retrieval with Contrastive Learning and Exponential Moving Average Distillation
por: Xiao, Ling, et al.
Publicado: (2026) -
Iterative Self-Improvement of Vision Language Models for Image Scoring and Self-Explanation
por: Tanji, Naoto, et al.
Publicado: (2025) -
Language-Guided Graph Representation Learning for Video Summarization
por: Li, Wenrui, et al.
Publicado: (2025) -
Language-guided Detection and Mitigation of Unknown Dataset Bias
por: Zhao, Zaiying, et al.
Publicado: (2024)