SCoRD: Subject-Conditional Relation Detection with Text-Augmented Data
Fuente:
arXiv
Saved in:
| Main Authors: | Yang, Ziyan, Kafle, Kushal, Lin, Zhe, Cohen, Scott, Ding, Zhihong, Ordonez, Vicente |
|---|---|
| Format: | Preprint |
| Published: |
2023
|
| Subjects: | |
| Online Access: | |
| Tags: |
Add Tag
No Tags, Be the first to tag this record!
|
Similar Items
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022)
by: Yang, Ziyan, et al.
Published: (2022)
Improving Large Vision and Language Models by Learning from a Panel of Peers
by: Hernandez, Jefferson, et al.
Published: (2025)
by: Hernandez, Jefferson, et al.
Published: (2025)
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025)
by: Koo, Jaywon, et al.
Published: (2025)
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
by: Hua, Hang, et al.
Published: (2024)
by: Hua, Hang, et al.
Published: (2024)
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
by: Slyman, Eric, et al.
Published: (2024)
by: Slyman, Eric, et al.
Published: (2024)
Plot'n Polish: Zero-shot Story Visualization and Disentangled Editing with Text-to-Image Diffusion Models
by: Akdemir, Kiymet, et al.
Published: (2025)
by: Akdemir, Kiymet, et al.
Published: (2025)
Visual Grounding Methods for VQA are Working for the Wrong Reasons!
by: Shrestha, Robik, et al.
Published: (2020)
by: Shrestha, Robik, et al.
Published: (2020)
Are Bias Mitigation Techniques for Deep Learning Effective?
by: Shrestha, Robik, et al.
Published: (2021)
by: Shrestha, Robik, et al.
Published: (2021)
Latent Feature-Guided Diffusion Models for Shadow Removal
by: Mei, Kangfu, et al.
Published: (2023)
by: Mei, Kangfu, et al.
Published: (2023)
Calibrating MLLM-as-a-judge via Multimodal Bayesian Prompt Ensembles
by: Slyman, Eric, et al.
Published: (2025)
by: Slyman, Eric, et al.
Published: (2025)
MetaShadow: Object-Centered Shadow Detection, Removal, and Synthesis
by: Wang, Tianyu, et al.
Published: (2024)
by: Wang, Tianyu, et al.
Published: (2024)
EntityBench: Towards Entity-Consistent Long-Range Multi-Shot Video Generation
by: He, Ruozhen, et al.
Published: (2026)
by: He, Ruozhen, et al.
Published: (2026)
CLIP-RD: Relative Distillation for Efficient CLIP Knowledge Distillation
by: Chung, Jeannie, et al.
Published: (2026)
by: Chung, Jeannie, et al.
Published: (2026)
PropTest: Automatic Property Testing for Improved Visual Programming
by: Koo, Jaywon, et al.
Published: (2024)
by: Koo, Jaywon, et al.
Published: (2024)
NoiseShift: Resolution-Aware Noise Recalibration for Better Low-Resolution Image Generation
by: He, Ruozhen, et al.
Published: (2025)
by: He, Ruozhen, et al.
Published: (2025)
D-SCo: Dual-Stream Conditional Diffusion for Monocular Hand-Held Object Reconstruction
by: Fu, Bowen, et al.
Published: (2023)
by: Fu, Bowen, et al.
Published: (2023)
Seeing Through Words: Controlling Visual Retrieval Quality with Language Models
by: Lu, Jianglin, et al.
Published: (2026)
by: Lu, Jianglin, et al.
Published: (2026)
Anatomy-Aware Conditional Image-Text Retrieval
by: Zheng, Meng, et al.
Published: (2025)
by: Zheng, Meng, et al.
Published: (2025)
Learning from Synthetic Data for Visual Grounding
by: He, Ruozhen, et al.
Published: (2024)
by: He, Ruozhen, et al.
Published: (2024)
RetouchIQ: MLLM Agents for Instruction-Based Image Retouching with Generalist Reward
by: Wu, Qiucheng, et al.
Published: (2026)
by: Wu, Qiucheng, et al.
Published: (2026)
SCoDA: Self-supervised Continual Domain Adaptation
by: Agrawal, Chirayu, et al.
Published: (2025)
by: Agrawal, Chirayu, et al.
Published: (2025)
They're All Doctors: Synthesizing Diverse Counterfactuals to Mitigate Associative Bias
by: Magid, Salma Abdel, et al.
Published: (2024)
by: Magid, Salma Abdel, et al.
Published: (2024)
TD-RD: A Top-Down Benchmark with Real-Time Framework for Road Damage Detection
by: Xiao, Xi, et al.
Published: (2025)
by: Xiao, Xi, et al.
Published: (2025)
OmniVCus: Feedforward Subject-driven Video Customization with Multimodal Control Conditions
by: Cai, Yuanhao, et al.
Published: (2025)
by: Cai, Yuanhao, et al.
Published: (2025)
Rethinking Global Text Conditioning in Diffusion Transformers
by: Starodubcev, Nikita, et al.
Published: (2026)
by: Starodubcev, Nikita, et al.
Published: (2026)
RD-VIO: Robust Visual-Inertial Odometry for Mobile Augmented Reality in Dynamic Environments
by: Li, Jinyu, et al.
Published: (2023)
by: Li, Jinyu, et al.
Published: (2023)
CompleteMe: Reference-based Human Image Completion
by: Tsai, Yu-Ju, et al.
Published: (2025)
by: Tsai, Yu-Ju, et al.
Published: (2025)
Controllable and Efficient Multi-Class Pathology Nuclei Data Augmentation using Text-Conditioned Diffusion Models
by: Oh, Hyun-Jic, et al.
Published: (2024)
by: Oh, Hyun-Jic, et al.
Published: (2024)
RetriBooru: Leakage-Free Retrieval of Conditions from Reference Images for Subject-Driven Generation
by: Tang, Haoran, et al.
Published: (2023)
by: Tang, Haoran, et al.
Published: (2023)
SCoFT: Self-Contrastive Fine-Tuning for Equitable Image Generation
by: Liu, Zhixuan, et al.
Published: (2024)
by: Liu, Zhixuan, et al.
Published: (2024)
SCoRe: Submodular Combinatorial Representation Learning
by: Majee, Anay, et al.
Published: (2023)
by: Majee, Anay, et al.
Published: (2023)
Group Relative Augmentation for Data Efficient Action Detection
by: Patel, Deep Anil, et al.
Published: (2025)
by: Patel, Deep Anil, et al.
Published: (2025)
On the Importance of Conditioning for Privacy-Preserving Data Augmentation
by: Lorenz, Julian, et al.
Published: (2025)
by: Lorenz, Julian, et al.
Published: (2025)
Data Augmentation for Text-based Person Retrieval Using Large Language Models
by: Li, Zheng, et al.
Published: (2024)
by: Li, Zheng, et al.
Published: (2024)
RAID: Retrieval-Augmented Anomaly Detection
by: Cai, Mingxiu, et al.
Published: (2026)
by: Cai, Mingxiu, et al.
Published: (2026)
More Than the Final Answer: Improving Visual Extraction and Logical Consistency in Vision-Language Models
by: Just, Hoang Anh, et al.
Published: (2025)
by: Just, Hoang Anh, et al.
Published: (2025)
SCoCCA: Multi-modal Sparse Concept Decomposition via Canonical Correlation Analysis
by: Gordon, Ehud, et al.
Published: (2026)
by: Gordon, Ehud, et al.
Published: (2026)
A-SCoRe: Attention-based Scene Coordinate Regression for wide-ranging scenarios
by: Bui, Huy-Hoang, et al.
Published: (2025)
by: Bui, Huy-Hoang, et al.
Published: (2025)
SCoRe: Clean Image Generation from Diffusion Models Trained on Noisy Images
by: Matsuzaki, Yuta, et al.
Published: (2026)
by: Matsuzaki, Yuta, et al.
Published: (2026)
Data Augmentation via Latent Diffusion Models for Detecting Smell-Related Objects in Historical Artworks
by: Sheta, Ahmed, et al.
Published: (2025)
by: Sheta, Ahmed, et al.
Published: (2025)
Similar Items
-
Improving Visual Grounding by Encouraging Consistent Gradient-based Explanations
by: Yang, Ziyan, et al.
Published: (2022) -
Improving Large Vision and Language Models by Learning from a Panel of Peers
by: Hernandez, Jefferson, et al.
Published: (2025) -
Evaluating Text-to-Image and Text-to-Video Synthesis with a Conditional Fréchet Distance
by: Koo, Jaywon, et al.
Published: (2025) -
FINEMATCH: Aspect-based Fine-grained Image and Text Mismatch Detection and Correction
by: Hua, Hang, et al.
Published: (2024) -
FairDeDup: Detecting and Mitigating Vision-Language Fairness Disparities in Semantic Dataset Deduplication
by: Slyman, Eric, et al.
Published: (2024)