GameTileNet: A Semantic Dataset for Low-Resolution Game Art in Procedural Content Generation
Fuente:
arXiv
Salvato in:
| Autori principali: | Chen, Yi-Chun, Jhala, Arnav |
|---|---|
| Natura: | Preprint |
| Pubblicazione: |
2025
|
| Soggetti: | |
| Accesso online: | |
| Tags: |
Aggiungi Tag
Nessun Tag, puoi essere il primo ad aggiungerne!!
|
Documenti analoghi
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
di: Chen, Yi-Chun, et al.
Pubblicazione: (2025)
di: Chen, Yi-Chun, et al.
Pubblicazione: (2025)
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
di: Lv, Zheqi, et al.
Pubblicazione: (2025)
di: Lv, Zheqi, et al.
Pubblicazione: (2025)
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
di: Wang, Junjie, et al.
Pubblicazione: (2024)
di: Wang, Junjie, et al.
Pubblicazione: (2024)
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
di: Zhang, Junzheng, et al.
Pubblicazione: (2024)
di: Zhang, Junzheng, et al.
Pubblicazione: (2024)
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
di: Chen, Yi-Chun
Pubblicazione: (2025)
di: Chen, Yi-Chun
Pubblicazione: (2025)
RS5M and GeoRSCLIP: A Large Scale Vision-Language Dataset and A Large Vision-Language Model for Remote Sensing
di: Zhang, Zilun, et al.
Pubblicazione: (2023)
di: Zhang, Zilun, et al.
Pubblicazione: (2023)
Vision as a Dialect: Unifying Visual Understanding and Generation via Text-Aligned Representations
di: Han, Jiaming, et al.
Pubblicazione: (2025)
di: Han, Jiaming, et al.
Pubblicazione: (2025)
VCap: Hypergeometric Rewards for Weak-to-Strong Visual Captioning
di: Lu, Xingyu, et al.
Pubblicazione: (2026)
di: Lu, Xingyu, et al.
Pubblicazione: (2026)
A Pilot Study on Curator-Guided Multilingual Art Description for Blind and Low-Vision Audiences with Small Vision-Language Models
di: Tsangko, Iosif, et al.
Pubblicazione: (2026)
di: Tsangko, Iosif, et al.
Pubblicazione: (2026)
Discriminative Probing and Tuning for Text-to-Image Generation
di: Qu, Leigang, et al.
Pubblicazione: (2024)
di: Qu, Leigang, et al.
Pubblicazione: (2024)
Order Is Not Layout: Order-to-Space Bias in Image Generation
di: Zhang, Yongkang, et al.
Pubblicazione: (2026)
di: Zhang, Yongkang, et al.
Pubblicazione: (2026)
Efficient Low-Resolution Face Recognition via Bridge Distillation
di: Ge, Shiming, et al.
Pubblicazione: (2024)
di: Ge, Shiming, et al.
Pubblicazione: (2024)
Wiki-LLaVA: Hierarchical Retrieval-Augmented Generation for Multimodal LLMs
di: Caffagni, Davide, et al.
Pubblicazione: (2024)
di: Caffagni, Davide, et al.
Pubblicazione: (2024)
ReAG: Reasoning-Augmented Generation for Knowledge-based Visual Question Answering
di: Compagnoni, Alberto, et al.
Pubblicazione: (2025)
di: Compagnoni, Alberto, et al.
Pubblicazione: (2025)
CalliffusionV2: Personalized Natural Calligraphy Generation with Flexible Multi-modal Control
di: Liao, Qisheng, et al.
Pubblicazione: (2024)
di: Liao, Qisheng, et al.
Pubblicazione: (2024)
CounterVid: Counterfactual Video Generation for Mitigating Action and Temporal Hallucinations in Video-Language Models
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
di: Poppi, Tobia, et al.
Pubblicazione: (2026)
LongLLaVA: Scaling Multi-modal LLMs to 1000 Images Efficiently via a Hybrid Architecture
di: Wang, Xidong, et al.
Pubblicazione: (2024)
di: Wang, Xidong, et al.
Pubblicazione: (2024)
Multimodal LLMs Can Reason about Aesthetics in Zero-Shot
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
di: Jiang, Ruixiang, et al.
Pubblicazione: (2025)
ADS-Edit: A Multimodal Knowledge Editing Dataset for Autonomous Driving Systems
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
di: Wang, Chenxi, et al.
Pubblicazione: (2025)
Self-Adaptive Sampling for Efficient Video Question-Answering on Image--Text Models
di: Han, Wei, et al.
Pubblicazione: (2023)
di: Han, Wei, et al.
Pubblicazione: (2023)
LLMs Meet Multimodal Generation and Editing: A Survey
di: He, Yingqing, et al.
Pubblicazione: (2024)
di: He, Yingqing, et al.
Pubblicazione: (2024)
VIEScore: Towards Explainable Metrics for Conditional Image Synthesis Evaluation
di: Ku, Max, et al.
Pubblicazione: (2023)
di: Ku, Max, et al.
Pubblicazione: (2023)
CLCR: Cross-Level Semantic Collaborative Representation for Multimodal Learning
di: Meng, Chunlei, et al.
Pubblicazione: (2026)
di: Meng, Chunlei, et al.
Pubblicazione: (2026)
DyRoNet: Dynamic Routing and Low-Rank Adapters for Autonomous Driving Streaming Perception
di: Huang, Xiang, et al.
Pubblicazione: (2024)
di: Huang, Xiang, et al.
Pubblicazione: (2024)
Both Text and Images Leaked! A Systematic Analysis of Data Contamination in Multimodal LLM
di: Song, Dingjie, et al.
Pubblicazione: (2024)
di: Song, Dingjie, et al.
Pubblicazione: (2024)
Less is More: A Simple yet Effective Token Reduction Method for Efficient Multi-modal LLMs
di: Song, Dingjie, et al.
Pubblicazione: (2024)
di: Song, Dingjie, et al.
Pubblicazione: (2024)
TheoremExplainAgent: Towards Video-based Multimodal Explanations for LLM Theorem Understanding
di: Ku, Max, et al.
Pubblicazione: (2025)
di: Ku, Max, et al.
Pubblicazione: (2025)
POLYCHARTQA: Benchmarking Large Vision-Language Models with Multilingual Chart Question Answering
di: Xu, Yichen, et al.
Pubblicazione: (2025)
di: Xu, Yichen, et al.
Pubblicazione: (2025)
Analyzing Diffusion and Autoregressive Vision Language Models in Multimodal Embedding Space
di: Wang, Zihang, et al.
Pubblicazione: (2026)
di: Wang, Zihang, et al.
Pubblicazione: (2026)
On Semiotic-Grounded Interpretive Evaluation of Generative Art
di: Jiang, Ruixiang, et al.
Pubblicazione: (2026)
di: Jiang, Ruixiang, et al.
Pubblicazione: (2026)
NExT-OMNI: Towards Any-to-Any Omnimodal Foundation Models with Discrete Flow Matching
di: Luo, Run, et al.
Pubblicazione: (2025)
di: Luo, Run, et al.
Pubblicazione: (2025)
Prompt-OT: An Optimal Transport Regularization Paradigm for Knowledge Preservation in Vision-Language Model Adaptation
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
di: Chen, Xiwen, et al.
Pubblicazione: (2025)
Official-NV: An LLM-Generated News Video Dataset for Multimodal Fake News Detection
di: Wang, Yihao, et al.
Pubblicazione: (2024)
di: Wang, Yihao, et al.
Pubblicazione: (2024)
Enhancing Environmental Monitoring through Multispectral Imaging: The WasteMS Dataset for Semantic Segmentation of Lakeside Waste
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
di: Zhu, Qinfeng, et al.
Pubblicazione: (2024)
Dual-Modal Attention-Enhanced Text-Video Retrieval with Triplet Partial Margin Contrastive Learning
di: Jiang, Chen, et al.
Pubblicazione: (2023)
di: Jiang, Chen, et al.
Pubblicazione: (2023)
Pegasus-v1 Technical Report
di: Jung, Raehyuk, et al.
Pubblicazione: (2024)
di: Jung, Raehyuk, et al.
Pubblicazione: (2024)
Bernini: Latent Semantic Planning for Video Diffusion
di: Bernini Team, et al.
Pubblicazione: (2026)
di: Bernini Team, et al.
Pubblicazione: (2026)
Harnessing Self-Supervised Features for Art Classification
di: Melis, Federico, et al.
Pubblicazione: (2026)
di: Melis, Federico, et al.
Pubblicazione: (2026)
The CASTLE 2024 Dataset: Advancing the Art of Multimodal Understanding
di: Rossetto, Luca, et al.
Pubblicazione: (2025)
di: Rossetto, Luca, et al.
Pubblicazione: (2025)
Heracles: A Hybrid SSM-Transformer Model for High-Resolution Image and Time-Series Analysis
di: Patro, Badri N., et al.
Pubblicazione: (2024)
di: Patro, Badri N., et al.
Pubblicazione: (2024)
Documenti analoghi
-
Narrative-to-Scene Generation: An LLM-Driven Pipeline for 2D Game Environments
di: Chen, Yi-Chun, et al.
Pubblicazione: (2025) -
Multimodal LLM-Guided Semantic Correction in Text-to-Image Diffusion
di: Lv, Zheqi, et al.
Pubblicazione: (2025) -
PIN: A Knowledge-Intensive Dataset for Paired and Interleaved Multimodal Documents
di: Wang, Junjie, et al.
Pubblicazione: (2024) -
Distilling Generative-Discriminative Representations for Very Low-Resolution Face Recognition
di: Zhang, Junzheng, et al.
Pubblicazione: (2024) -
Hierarchical Knowledge Graphs for Story Understanding in Visual Narratives
di: Chen, Yi-Chun
Pubblicazione: (2025)