Gespeichert in:
| Hauptverfasser: | Dong, Ziyi, Zhang, Yurui, Li, Changmao, Golding, Naomi Rue, Long, Qing |
|---|---|
| Format: | Preprint |
| Veröffentlicht: |
2025
|
| Schlagworte: | |
| Online-Zugang: | https://arxiv.org/abs/2510.07951 |
| Tags: |
Tag hinzufügen
Keine Tags, Fügen Sie den ersten Tag hinzu!
|
Ähnliche Einträge
FurniScene: A Large-scale 3D Room Dataset with Intricate Furnishing Scenes
von: Zhang, Genghao, et al.
Veröffentlicht: (2024)
von: Zhang, Genghao, et al.
Veröffentlicht: (2024)
WorldAfford: Affordance Grounding based on Natural Language Instructions
von: Chen, Changmao, et al.
Veröffentlicht: (2024)
von: Chen, Changmao, et al.
Veröffentlicht: (2024)
SceneX: Procedural Controllable Large-scale Scene Generation
von: Zhou, Mengqi, et al.
Veröffentlicht: (2024)
von: Zhou, Mengqi, et al.
Veröffentlicht: (2024)
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
von: Simonyan, Aleksandr, et al.
Veröffentlicht: (2026)
von: Simonyan, Aleksandr, et al.
Veröffentlicht: (2026)
APISR: Anime Production Inspired Real-World Anime Super-Resolution
von: Wang, Boyang, et al.
Veröffentlicht: (2024)
von: Wang, Boyang, et al.
Veröffentlicht: (2024)
LinkTo-Anime: A 2D Animation Optical Flow Dataset from 3D Model Rendering
von: Feng, Xiaoyi, et al.
Veröffentlicht: (2025)
von: Feng, Xiaoyi, et al.
Veröffentlicht: (2025)
GraspClutter6D: A Large-scale Real-world Dataset for Robust Perception and Grasping in Cluttered Scenes
von: Back, Seunghyeok, et al.
Veröffentlicht: (2025)
von: Back, Seunghyeok, et al.
Veröffentlicht: (2025)
EventSTR: A Benchmark Dataset and Baselines for Event Stream based Scene Text Recognition
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
von: Wang, Xiao, et al.
Veröffentlicht: (2025)
Enhanced Anime Image Generation Using USE-CMHSA-GAN
von: Lu, J.
Veröffentlicht: (2024)
von: Lu, J.
Veröffentlicht: (2024)
FocusDD: Real-World Scene Infusion for Robust Dataset Distillation
von: Hu, Youbing, et al.
Veröffentlicht: (2025)
von: Hu, Youbing, et al.
Veröffentlicht: (2025)
Region Prompt Tuning: Fine-grained Scene Text Detection Utilizing Region Text Prompt
von: Lin, Xingtao, et al.
Veröffentlicht: (2024)
von: Lin, Xingtao, et al.
Veröffentlicht: (2024)
CURVE: Learning Causality-Inspired Invariant Representations for Robust Scene Understanding via Uncertainty-Guided Regularization
von: Liang, Yue, et al.
Veröffentlicht: (2026)
von: Liang, Yue, et al.
Veröffentlicht: (2026)
A Large-Scale Multimodal Dataset and Benchmarks for Human Activity Scene Understanding and Reasoning
von: Jiang, Siyang, et al.
Veröffentlicht: (2025)
von: Jiang, Siyang, et al.
Veröffentlicht: (2025)
GauU-Scene: A Scene Reconstruction Benchmark on Large Scale 3D Reconstruction Dataset Using Gaussian Splatting
von: Xiong, Butian, et al.
Veröffentlicht: (2024)
von: Xiong, Butian, et al.
Veröffentlicht: (2024)
Bharat Scene Text: A Novel Comprehensive Dataset and Benchmark for Indian Language Scene Text Understanding
von: De, Anik, et al.
Veröffentlicht: (2025)
von: De, Anik, et al.
Veröffentlicht: (2025)
DreamText: High Fidelity Scene Text Synthesis
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
von: Wang, Yibin, et al.
Veröffentlicht: (2024)
RoleMotion: A Large-Scale Dataset towards Robust Scene-Specific Role-Playing Motion Synthesis with Fine-grained Descriptions
von: Peng, Junran, et al.
Veröffentlicht: (2025)
von: Peng, Junran, et al.
Veröffentlicht: (2025)
STAR: A First-Ever Dataset and A Large-Scale Benchmark for Scene Graph Generation in Large-Size Satellite Imagery
von: Li, Yansheng, et al.
Veröffentlicht: (2024)
von: Li, Yansheng, et al.
Veröffentlicht: (2024)
TextMamba: Scene Text Detector with Mamba
von: Zhao, Qiyan, et al.
Veröffentlicht: (2025)
von: Zhao, Qiyan, et al.
Veröffentlicht: (2025)
3D-SceneDreamer: Text-Driven 3D-Consistent Scene Generation
von: Zhang, Frank, et al.
Veröffentlicht: (2024)
von: Zhang, Frank, et al.
Veröffentlicht: (2024)
Can Vision-Language Models be a Good Guesser? Exploring VLMs for Times and Location Reasoning
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
von: Zhang, Gengyuan, et al.
Veröffentlicht: (2023)
From Easy to Hard: Learning Curricular Shape-aware Features for Robust Panoptic Scene Graph Generation
von: Shi, Hanrong, et al.
Veröffentlicht: (2024)
von: Shi, Hanrong, et al.
Veröffentlicht: (2024)
Text-Scene: A Scene-to-Language Parsing Framework for 3D Scene Understanding
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
von: Li, Haoyuan, et al.
Veröffentlicht: (2025)
MPerS: Dynamic MLLM MixExperts Perception-Guided Remote Sensing Scene Segmentation
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
von: Wang, Ziyi, et al.
Veröffentlicht: (2026)
WAS: Dataset and Methods for Artistic Text Segmentation
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
von: Xie, Xudong, et al.
Veröffentlicht: (2024)
AUG: A New Dataset and An Efficient Model for Aerial Image Urban Scene Graph Generation
von: Li, Yansheng, et al.
Veröffentlicht: (2024)
von: Li, Yansheng, et al.
Veröffentlicht: (2024)
SSIMBaD: Sigma Scaling with SSIM-Guided Balanced Diffusion for AnimeFace Colorization
von: Seo, Junpyo, et al.
Veröffentlicht: (2025)
von: Seo, Junpyo, et al.
Veröffentlicht: (2025)
OmniScience: A Large-scale Multi-modal Dataset for Scientific Image Understanding
von: Tao, Haoyi, et al.
Veröffentlicht: (2026)
von: Tao, Haoyi, et al.
Veröffentlicht: (2026)
MPT: A Large-scale Multi-Phytoplankton Tracking Benchmark
von: Yu, Yang, et al.
Veröffentlicht: (2024)
von: Yu, Yang, et al.
Veröffentlicht: (2024)
InstructOCR: Instruction Boosting Scene Text Spotting
von: Duan, Chen, et al.
Veröffentlicht: (2024)
von: Duan, Chen, et al.
Veröffentlicht: (2024)
3DCoMPaT$^{++}$: An improved Large-scale 3D Vision Dataset for Compositional Recognition
von: Slim, Habib, et al.
Veröffentlicht: (2023)
von: Slim, Habib, et al.
Veröffentlicht: (2023)
RadiomicsRetrieval: A Customizable Framework for Medical Image Retrieval Using Radiomics Features
von: Na, Inye, et al.
Veröffentlicht: (2025)
von: Na, Inye, et al.
Veröffentlicht: (2025)
SurgMLLMBench: A Multimodal Large Language Model Benchmark Dataset for Surgical Scene Understanding
von: Choi, Tae-Min, et al.
Veröffentlicht: (2025)
von: Choi, Tae-Min, et al.
Veröffentlicht: (2025)
Finding Outliers in a Haystack: Anomaly Detection for Large Pointcloud Scenes
von: Faulkner, Ryan, et al.
Veröffentlicht: (2025)
von: Faulkner, Ryan, et al.
Veröffentlicht: (2025)
RU-AI: A Large Multimodal Dataset for Machine-Generated Content Detection
von: Huang, Liting, et al.
Veröffentlicht: (2024)
von: Huang, Liting, et al.
Veröffentlicht: (2024)
ParkingScenes: A Structured Dataset for End-to-End Autonomous Parking in Simulation Scenes
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
von: Chen, Haonan, et al.
Veröffentlicht: (2026)
Adversarial Training with OCR Modality Perturbation for Scene-Text Visual Question Answering
von: Shen, Zhixuan, et al.
Veröffentlicht: (2024)
von: Shen, Zhixuan, et al.
Veröffentlicht: (2024)
15M Multimodal Facial Image-Text Dataset
von: Dai, Dawei, et al.
Veröffentlicht: (2024)
von: Dai, Dawei, et al.
Veröffentlicht: (2024)
M2-Encoder: Advancing Bilingual Image-Text Understanding by Large-scale Efficient Pretraining
von: Guo, Qingpei, et al.
Veröffentlicht: (2024)
von: Guo, Qingpei, et al.
Veröffentlicht: (2024)
DecepGPT: Schema-Driven Deception Detection with Multicultural Datasets and Robust Multimodal Learning
von: Huang, Jiajian, et al.
Veröffentlicht: (2026)
von: Huang, Jiajian, et al.
Veröffentlicht: (2026)
Ähnliche Einträge
-
FurniScene: A Large-scale 3D Room Dataset with Intricate Furnishing Scenes
von: Zhang, Genghao, et al.
Veröffentlicht: (2024) -
WorldAfford: Affordance Grounding based on Natural Language Instructions
von: Chen, Changmao, et al.
Veröffentlicht: (2024) -
SceneX: Procedural Controllable Large-scale Scene Generation
von: Zhou, Mengqi, et al.
Veröffentlicht: (2024) -
StyleText: A Large-Scale Dataset and Benchmark for Stylized Scene Text Inpainting
von: Simonyan, Aleksandr, et al.
Veröffentlicht: (2026) -
APISR: Anime Production Inspired Real-World Anime Super-Resolution
von: Wang, Boyang, et al.
Veröffentlicht: (2024)