Exploring Semantic Perturbations on Grover
Fuente:
arXiv
Gespeichert in:
| Hauptverfasser: | Ji, Ziqing, Kulkarni, Pranav, Neskovic, Marko, Nolan, Kevin, Xu, Yan |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2023
|
| Schlagworte: | |
| Online-Zugang: | |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
Semantic Prompt Learning for Weakly-Supervised Semantic Segmentation
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024)
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023)
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
von: Xu, Xiao, et al.
Veröffentlicht: (2024)
Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation
von: Banerjee, Oishi, et al.
Veröffentlicht: (2024)
von: Banerjee, Oishi, et al.
Veröffentlicht: (2024)
BiCLIP: Domain Canonicalization via Structured Geometric Transformation
von: Mantini, Pranav, et al.
Veröffentlicht: (2026)
von: Mantini, Pranav, et al.
Veröffentlicht: (2026)
Sensitivity of Generative VLMs to Semantically and Lexically Altered Prompts
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
von: Dumpala, Sri Harsha, et al.
Veröffentlicht: (2024)
MoTe: Learning Motion-Text Diffusion Model for Multiple Generation Tasks
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
von: Wu, Yiming, et al.
Veröffentlicht: (2024)
Semantic Token Reweighting for Interpretable and Controllable Text Embeddings in CLIP
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
von: Kim, Eunji, et al.
Veröffentlicht: (2024)
SemIRNet: A Semantic Irony Recognition Network for Multimodal Sarcasm Detection
von: Zhou, Jingxuan, et al.
Veröffentlicht: (2025)
von: Zhou, Jingxuan, et al.
Veröffentlicht: (2025)
Exploring CLIP's Dense Knowledge for Weakly Supervised Semantic Segmentation
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
von: Yang, Zhiwei, et al.
Veröffentlicht: (2025)
VisualRWKV: Exploring Recurrent Neural Networks for Visual Language Models
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
von: Hou, Haowen, et al.
Veröffentlicht: (2024)
MM-Spatial: Exploring 3D Spatial Understanding in Multimodal LLMs
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)
von: Daxberger, Erik, et al.
Veröffentlicht: (2025)
mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
von: Ye, Qinghao, et al.
Veröffentlicht: (2023)
Training-Free Generation of Diverse and High-Fidelity Images via Prompt Semantic Space Optimization
von: Meng, Debin, et al.
Veröffentlicht: (2025)
von: Meng, Debin, et al.
Veröffentlicht: (2025)
Semantic-Clipping: Efficient Vision-Language Modeling with Semantic-Guidedd Visual Selection
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
von: Li, Bangzheng, et al.
Veröffentlicht: (2025)
Set-CLIP: Exploring Aligned Semantic From Low-Alignment Multimodal Data Through A Distribution View
von: Song, Zijia, et al.
Veröffentlicht: (2024)
von: Song, Zijia, et al.
Veröffentlicht: (2024)
A Semantic Space is Worth 256 Language Descriptions: Make Stronger Segmentation Models with Descriptive Properties
von: Xiao, Junfei, et al.
Veröffentlicht: (2023)
von: Xiao, Junfei, et al.
Veröffentlicht: (2023)
Gloss2Text: Sign Language Gloss translation using LLMs and Semantically Aware Label Smoothing
von: Fayyazsanavi, Pooya, et al.
Veröffentlicht: (2024)
von: Fayyazsanavi, Pooya, et al.
Veröffentlicht: (2024)
Exploring Attention Mechanisms in Integration of Multi-Modal Information for Sign Language Recognition and Translation
von: Hakim, Zaber Ibn Abdul, et al.
Veröffentlicht: (2023)
von: Hakim, Zaber Ibn Abdul, et al.
Veröffentlicht: (2023)
Exploring How Generative MLLMs Perceive More Than CLIP with the Same Vision Encoder
von: Li, Siting, et al.
Veröffentlicht: (2024)
von: Li, Siting, et al.
Veröffentlicht: (2024)
Towards Accessible Learning: Deep Learning-Based Potential Dysgraphia Detection and OCR for Potentially Dysgraphic Handwriting
von: D, Vydeki, et al.
Veröffentlicht: (2024)
von: D, Vydeki, et al.
Veröffentlicht: (2024)
Exploring the Limits of Zero Shot Vision Language Models for Hate Meme Detection: The Vulnerabilities and their Interpretations
von: Rizwan, Naquee, et al.
Veröffentlicht: (2024)
von: Rizwan, Naquee, et al.
Veröffentlicht: (2024)
Scaling Inference-Time Search with Vision Value Model for Improved Visual Comprehension
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
von: Wang, Xiyao, et al.
Veröffentlicht: (2024)
Prompt as Free Lunch: Enhancing Diversity in Source-Free Cross-domain Few-shot Learning through Semantic-Guided Prompting
von: Zhuo, Linhai, et al.
Veröffentlicht: (2024)
von: Zhuo, Linhai, et al.
Veröffentlicht: (2024)
X-Mark: Saliency-Guided Robust Dataset Ownership Verification for Medical Imaging
von: Kulkarni, Pranav, et al.
Veröffentlicht: (2026)
von: Kulkarni, Pranav, et al.
Veröffentlicht: (2026)
S-GRPO: Unified Post-Training for Large Vision-Language Models
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
von: Yan, Yuming, et al.
Veröffentlicht: (2026)
SOLO: A Single Transformer for Scalable Vision-Language Modeling
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
von: Chen, Yangyi, et al.
Veröffentlicht: (2024)
Prioritizing Image-Related Tokens Enhances Vision-Language Pre-Training
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
von: Chen, Yangyi, et al.
Veröffentlicht: (2025)
Measuring and Improving Chain-of-Thought Reasoning in Vision-Language Models
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
DRESS: Instructing Large Vision-Language Models to Align and Interact with Humans via Natural Language Feedback
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
von: Chen, Yangyi, et al.
Veröffentlicht: (2023)
LLM as a Complementary Optimizer to Gradient Descent: A Case Study in Prompt Tuning
von: Guo, Zixian, et al.
Veröffentlicht: (2024)
von: Guo, Zixian, et al.
Veröffentlicht: (2024)
Advanced Multimodal Deep Learning Architecture for Image-Text Matching
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
von: Wang, Jinyin, et al.
Veröffentlicht: (2024)
Sparrow: Data-Efficient Video-LLM with Text-to-Image Augmentation
von: Yin, Shukang, et al.
Veröffentlicht: (2024)
von: Yin, Shukang, et al.
Veröffentlicht: (2024)
Scaling Up Temporal Domain Generalization via Temporal Experts Averaging
von: Liu, Aoming, et al.
Veröffentlicht: (2025)
von: Liu, Aoming, et al.
Veröffentlicht: (2025)
Gradient descent with generalized Newton's method
von: Bu, Zhiqi, et al.
Veröffentlicht: (2024)
von: Bu, Zhiqi, et al.
Veröffentlicht: (2024)
Towards Better Multi-head Attention via Channel-wise Sample Permutation
von: Yuan, Shen, et al.
Veröffentlicht: (2024)
von: Yuan, Shen, et al.
Veröffentlicht: (2024)
LiME: Lightweight Mixture of Experts for Efficient Multimodal Multi-task Learning
von: Kowsher, Md, et al.
Veröffentlicht: (2026)
von: Kowsher, Md, et al.
Veröffentlicht: (2026)
Verbalized Representation Learning for Interpretable Few-Shot Generalization
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
von: Yang, Cheng-Fu, et al.
Veröffentlicht: (2024)
Breaking through the learning plateaus of in-context learning in Transformer
von: Fu, Jingwen, et al.
Veröffentlicht: (2023)
von: Fu, Jingwen, et al.
Veröffentlicht: (2023)
Rethinking RL Scaling for Vision Language Models: A Transparent, From-Scratch Framework and Comprehensive Evaluation Scheme
von: Ma, Yan, et al.
Veröffentlicht: (2025)
von: Ma, Yan, et al.
Veröffentlicht: (2025)
Ähnliche Einträge
-
Semantic Prompt Learning for Weakly-Supervised Semantic Segmentation
von: Lin, Ci-Siang, et al.
Veröffentlicht: (2024) -
TiMix: Text-aware Image Mixing for Effective Vision-Language Pre-training
von: Jiang, Chaoya, et al.
Veröffentlicht: (2023) -
Exploring Multi-Grained Concept Annotations for Multimodal Large Language Models
von: Xu, Xiao, et al.
Veröffentlicht: (2024) -
Direct Preference Optimization for Suppressing Hallucinated Prior Exams in Radiology Report Generation
von: Banerjee, Oishi, et al.
Veröffentlicht: (2024) -
BiCLIP: Domain Canonicalization via Structured Geometric Transformation
von: Mantini, Pranav, et al.
Veröffentlicht: (2026)