Saved in:
| Main Authors: | Zhou, You, Wang, Jie |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | https://arxiv.org/abs/2311.01673 |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
by: Bucher, Martin JJ., et al.
Published: (2025)
by: Bucher, Martin JJ., et al.
Published: (2025)
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024)
by: Toker, Michael, et al.
Published: (2024)
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024)
by: Cao, Jingtao, et al.
Published: (2024)
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
by: Yang, Qi, et al.
Published: (2026)
by: Yang, Qi, et al.
Published: (2026)
Depthwise Separable Convolutions with Deep Residual Convolutions
by: Hasan, Md Arid, et al.
Published: (2024)
by: Hasan, Md Arid, et al.
Published: (2024)
R-Genie: Reasoning-Guided Generative Image Editing
by: Zhang, Dong, et al.
Published: (2025)
by: Zhang, Dong, et al.
Published: (2025)
Predicting 3D Rigid Body Dynamics with Deep Residual Network
by: Oketunji, Abiodun Finbarrs
Published: (2024)
by: Oketunji, Abiodun Finbarrs
Published: (2024)
Cinéaste: A Fine-grained Contextual Movie Question Answering Benchmark
by: Shah, Nisarg A., et al.
Published: (2025)
by: Shah, Nisarg A., et al.
Published: (2025)
VidNum-1.4K: A Comprehensive Benchmark for Video-based Numerical Reasoning
by: Cui, Shaoyang, et al.
Published: (2026)
by: Cui, Shaoyang, et al.
Published: (2026)
Enhanced Kalman with Adaptive Appearance Motion SORT for Grounded Generic Multiple Object Tracking
by: Anh, Duy Le Dinh, et al.
Published: (2024)
by: Anh, Duy Le Dinh, et al.
Published: (2024)
Multi-Reward GRPO for Stable and Prosodic Single-Codebook TTS LLMs at Scale
by: Zhong, Yicheng, et al.
Published: (2025)
by: Zhong, Yicheng, et al.
Published: (2025)
HalalBench: A Multilingual OCR Benchmark for Food Packaging Ingredient Extraction
by: Arief, Hasan
Published: (2026)
by: Arief, Hasan
Published: (2026)
Hyperspectral Images Efficient Spatial and Spectral non-Linear Model with Bidirectional Feature Learning
by: Yang, Judy X, et al.
Published: (2024)
by: Yang, Judy X, et al.
Published: (2024)
ReFACT: Updating Text-to-Image Models by Editing the Text Encoder
by: Arad, Dana, et al.
Published: (2023)
by: Arad, Dana, et al.
Published: (2023)
Unsupervised Band Selection Using Fused HSI and LiDAR Attention Integrating With Autoencoder
by: Yang, Judy X, et al.
Published: (2024)
by: Yang, Judy X, et al.
Published: (2024)
HSIMamba: Hyperpsectral Imaging Efficient Feature Learning with Bidirectional State Space for Classification
by: Yang, Judy X, et al.
Published: (2024)
by: Yang, Judy X, et al.
Published: (2024)
Class Incremental Learning with Task-Specific Batch Normalization and Out-of-Distribution Detection
by: Zhou, Zhiping, et al.
Published: (2024)
by: Zhou, Zhiping, et al.
Published: (2024)
Open-Set Supervised 3D Anomaly Detection: An Industrial Dataset and a Generalisable Framework for Unknown Defects
by: Liang, Hanzhe, et al.
Published: (2026)
by: Liang, Hanzhe, et al.
Published: (2026)
Text-to-Events: Synthetic Event Camera Streams from Conditional Text Input
by: Ott, Joachim, et al.
Published: (2024)
by: Ott, Joachim, et al.
Published: (2024)
Towards Explainable Fake Image Detection with Multi-Modal Large Language Models
by: Ji, Yikun, et al.
Published: (2025)
by: Ji, Yikun, et al.
Published: (2025)
Evaluation Before Generation: A Paradigm for Robust Multimodal Sentiment Analysis with Missing Modalities
by: Chen, Rongfei, et al.
Published: (2026)
by: Chen, Rongfei, et al.
Published: (2026)
Proximity QA: Unleashing the Power of Multi-Modal Large Language Models for Spatial Proximity Analysis
by: Li, Jianing, et al.
Published: (2024)
by: Li, Jianing, et al.
Published: (2024)
Only Whats Necessary: Pareto Optimal Data Minimization for Privacy Preserving Video Anomaly Detection
by: Aslam, Nazia, et al.
Published: (2026)
by: Aslam, Nazia, et al.
Published: (2026)
Analyze-Prompt-Reason: A Collaborative Agent-Based Framework for Multi-Image Vision-Language Reasoning
by: Vlachos, Angelos, et al.
Published: (2025)
by: Vlachos, Angelos, et al.
Published: (2025)
Safeguarding Vision-Language Models Against Patched Visual Prompt Injectors
by: Sun, Jiachen, et al.
Published: (2024)
by: Sun, Jiachen, et al.
Published: (2024)
Large Language Models for Biomedical Article Classification
by: Proboszcz, Jakub, et al.
Published: (2026)
by: Proboszcz, Jakub, et al.
Published: (2026)
GroundCap: A Visually Grounded Image Captioning Dataset
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
Relative Drawing Identification Complexity is Invariant to Modality in Vision-Language Models
by: Freitas, Diogo, et al.
Published: (2025)
by: Freitas, Diogo, et al.
Published: (2025)
More Than Meets the Eye: Measuring the Semiotic Gap in Vision-Language Models via Semantic Anchorage
by: He, Wei
Published: (2026)
by: He, Wei
Published: (2026)
StoryReasoning Dataset: Using Chain-of-Thought for Scene Understanding and Grounded Story Generation
by: Oliveira, Daniel A. P., et al.
Published: (2025)
by: Oliveira, Daniel A. P., et al.
Published: (2025)
Towards Blind and Low-Vision Accessibility of Lightweight VLMs and Custom LLM-Evals
by: Baghel, Shruti Singh, et al.
Published: (2025)
by: Baghel, Shruti Singh, et al.
Published: (2025)
The American Sign Language Knowledge Graph: Infusing ASL Models with Linguistic Knowledge
by: Kezar, Lee, et al.
Published: (2024)
by: Kezar, Lee, et al.
Published: (2024)
CCS: Clinical Consensus Selection for Radiology Report Generation
by: Zhang, Xi, et al.
Published: (2026)
by: Zhang, Xi, et al.
Published: (2026)
MyoSem: Aligning Electromyography to Natural-Language Action Semantics for Hand Action Understanding
by: Wang, Chiyue, et al.
Published: (2026)
by: Wang, Chiyue, et al.
Published: (2026)
InterChart: Benchmarking Visual Reasoning Across Decomposed and Distributed Chart Information
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2025)
by: Iyengar, Anirudh Iyengar Kaniyar Narayana, et al.
Published: (2025)
jina-clip-v2: Multilingual Multimodal Embeddings for Text and Images
by: Koukounas, Andreas, et al.
Published: (2024)
by: Koukounas, Andreas, et al.
Published: (2024)
Unveiling Glitches: A Deep Dive into Image Encoding Bugs within CLIP
by: Ranjan, Ayush, et al.
Published: (2024)
by: Ranjan, Ayush, et al.
Published: (2024)
CAMME: Adaptive Deepfake Image Detection with Multi-Modal Cross-Attention
by: Khan, Naseem, et al.
Published: (2025)
by: Khan, Naseem, et al.
Published: (2025)
Patchfinder: Leveraging Visual Language Models for Accurate Information Retrieval using Model Uncertainty
by: Colman, Roman, et al.
Published: (2024)
by: Colman, Roman, et al.
Published: (2024)
Enhancement Without Contrast: Stability-Aware Multicenter Machine Learning for Glioma MRI Imaging
by: Amiri, Sajad, et al.
Published: (2025)
by: Amiri, Sajad, et al.
Published: (2025)
Similar Items
-
ReSpace: Text-Driven Autoregressive 3D Indoor Scene Synthesis and Editing
by: Bucher, Martin JJ., et al.
Published: (2025) -
Diffusion Lens: Interpreting Text Encoders in Text-to-Image Pipelines
by: Toker, Michael, et al.
Published: (2024) -
VLEU: a Method for Automatic Evaluation for Generalizability of Text-to-Image Models
by: Cao, Jingtao, et al.
Published: (2024) -
BabelDOC: Better Layout-Preserving PDF Translation via Intermediate Representation
by: Yang, Qi, et al.
Published: (2026) -
Depthwise Separable Convolutions with Deep Residual Convolutions
by: Hasan, Md Arid, et al.
Published: (2024)