NaiLIA: Multimodal Nail Design Retrieval Based on Dense Intent Descriptions and Palette Queries
Fuente:
arXiv
Guardado en:
| Autores principales: | Amemiya, Kanon, Yashima, Daichi, Katsumata, Kei, Komatsu, Takumi, Korekata, Ryosuke, Otsuki, Seitaro, Sugiura, Komei |
|---|---|
| Formato: | Preprint |
| Publicado: |
2026
|
| Materias: | |
| Acceso en línea: | |
| Etiquetas: |
Agregar Etiqueta
Sin Etiquetas, Sea el primero en etiquetar este registro!
|
Ejemplares similares
Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning with Dense Labeling
por: Yashima, Daichi, et al.
Publicado: (2024)
por: Yashima, Daichi, et al.
Publicado: (2024)
Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement
por: Katsumata, Kei, et al.
Publicado: (2025)
por: Katsumata, Kei, et al.
Publicado: (2025)
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
por: Yashima, Daichi, et al.
Publicado: (2026)
por: Yashima, Daichi, et al.
Publicado: (2026)
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
por: Korekata, Ryosuke, et al.
Publicado: (2025)
por: Korekata, Ryosuke, et al.
Publicado: (2025)
Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations
por: Goko, Miyu, et al.
Publicado: (2024)
por: Goko, Miyu, et al.
Publicado: (2024)
ReMoRa: Multimodal Large Language Model based on Refined Motion Representation for Long-Video Understanding
por: Yashima, Daichi, et al.
Publicado: (2026)
por: Yashima, Daichi, et al.
Publicado: (2026)
DM2RM: Dual-Mode Multimodal Ranking for Target Objects and Receptacles Based on Open-Vocabulary Instructions
por: Korekata, Ryosuke, et al.
Publicado: (2024)
por: Korekata, Ryosuke, et al.
Publicado: (2024)
MLLM-as-a-Judge Exhibits Model Preference Bias
por: Koyama, Shuitsu, et al.
Publicado: (2026)
por: Koyama, Shuitsu, et al.
Publicado: (2026)
VELA: An LLM-Hybrid-as-a-Judge Approach for Evaluating Long Image Captions
por: Matsuda, Kazuki, et al.
Publicado: (2025)
por: Matsuda, Kazuki, et al.
Publicado: (2025)
LLM-Free Image Captioning Evaluation in Reference-Flexible Settings
por: Hirano, Shinnosuke, et al.
Publicado: (2025)
por: Hirano, Shinnosuke, et al.
Publicado: (2025)
Stitch4D: Sparse Multi-Location 4D Urban Reconstruction via Spatio-Temporal Interpolation
por: Kogure, Hina, et al.
Publicado: (2026)
por: Kogure, Hina, et al.
Publicado: (2026)
HiFlow: Tokenization-Free Scale-Wise Autoregressive Policy Learning via Flow Matching
por: Yashima, Daichi, et al.
Publicado: (2026)
por: Yashima, Daichi, et al.
Publicado: (2026)
Layer-Wise Relevance Propagation with Conservation Property for ResNet
por: Otsuki, Seitaro, et al.
Publicado: (2024)
por: Otsuki, Seitaro, et al.
Publicado: (2024)
AnoleVLA: Lightweight Vision-Language-Action Model with Deep State Space Models for Mobile Manipulation
por: Takagi, Yusuke, et al.
Publicado: (2026)
por: Takagi, Yusuke, et al.
Publicado: (2026)
Polos: Multimodal Metric Learning from Human Feedback for Image Captioning
por: Wada, Yuiga, et al.
Publicado: (2024)
por: Wada, Yuiga, et al.
Publicado: (2024)
GENNAV: Polygon Mask Generation for Generalized Referring Navigable Regions
por: Katsumata, Kei, et al.
Publicado: (2025)
por: Katsumata, Kei, et al.
Publicado: (2025)
Nearest Neighbor Future Captioning: Generating Descriptions for Possible Collisions in Object Placement Tasks
por: Komatsu, Takumi, et al.
Publicado: (2024)
por: Komatsu, Takumi, et al.
Publicado: (2024)
Pre-Manipulation Alignment Prediction with Parallel Deep State-Space and Transformer Models
por: Kambara, Motonari, et al.
Publicado: (2025)
por: Kambara, Motonari, et al.
Publicado: (2025)
Future Success Prediction in Open-Vocabulary Object Manipulation Tasks Based on End-Effector Trajectories
por: Kambara, Motonari, et al.
Publicado: (2024)
por: Kambara, Motonari, et al.
Publicado: (2024)
Deep Space Weather Model: Long-Range Solar Flare Prediction from Multi-Wavelength Images
por: Nagashima, Shunya, et al.
Publicado: (2025)
por: Nagashima, Shunya, et al.
Publicado: (2025)
ZINA: Multimodal Fine-grained Hallucination Detection and Editing
por: Wada, Yuiga, et al.
Publicado: (2025)
por: Wada, Yuiga, et al.
Publicado: (2025)
Object Segmentation from Open-Vocabulary Manipulation Instructions Based on Optimal Transport Polygon Matching with Multimodal Foundation Models
por: Nishimura, Takayuki, et al.
Publicado: (2024)
por: Nishimura, Takayuki, et al.
Publicado: (2024)
DENEB: A Hallucination-Robust Automatic Evaluation Metric for Image Captioning
por: Matsuda, Kazuki, et al.
Publicado: (2024)
por: Matsuda, Kazuki, et al.
Publicado: (2024)
FLARE-SSM: Deep State Space Models with Influence-Balanced Loss for 72-Hour Solar Flare Prediction
por: Takagi, Yusuke, et al.
Publicado: (2025)
por: Takagi, Yusuke, et al.
Publicado: (2025)
Co-Scale Cross-Attentional Transformer for Rearrangement Target Detection
por: Matsuo, Haruka, et al.
Publicado: (2024)
por: Matsuo, Haruka, et al.
Publicado: (2024)
Attention Lattice Adapter: Visual Explanation Generation for Visual Foundation Model
por: Hirano, Shinnosuke, et al.
Publicado: (2025)
por: Hirano, Shinnosuke, et al.
Publicado: (2025)
Cortical-SSM: A Deep State Space Model for EEG and ECoG Motor Imagery Decoding
por: Suzuki, Shuntaro, et al.
Publicado: (2025)
por: Suzuki, Shuntaro, et al.
Publicado: (2025)
Abundance of pteropods in the Aegean Sea during LIA07, LIA08, LIA09 and LIA10
por: Siokou-Frangou, Ioanna, et al.
Publicado: (2014)
por: Siokou-Frangou, Ioanna, et al.
Publicado: (2014)
A bordo del Nai'a. Buceando en Fidji
Publicado: (1998)
Publicado: (1998)
NaiAD: Initiate Data-Driven Research for LLM Advertising
por: Zhang, Yihang, et al.
Publicado: (2026)
por: Zhang, Yihang, et al.
Publicado: (2026)
Zooplankton community in Thi Nai lagoon in the period of 2001-2020
por: Nguyen, Tam Vinh
Publicado: (2020)
por: Nguyen, Tam Vinh
Publicado: (2020)
Leaving berlin / Joseph Kanon
por: Kanon, Joseph
Publicado: (2015)
por: Kanon, Joseph
Publicado: (2015)
Toward a holistic tophus assessment in gout clinical trials: What lies beyond tophus count and size?
por: Kanon Jatuworapruk
Publicado: (2024)
por: Kanon Jatuworapruk
Publicado: (2024)
3DFlowRenderer: One-shot Face Re-enactment via Dense 3D Facial Flow Estimation
por: Nijhawan, Siddharth, et al.
Publicado: (2024)
por: Nijhawan, Siddharth, et al.
Publicado: (2024)
QUIDS: Query Intent Description for Exploratory Search via Dual Space Modeling
por: Wang, Yumeng, et al.
Publicado: (2024)
por: Wang, Yumeng, et al.
Publicado: (2024)
Fixed Very‐Low‐Dose Oral Immunotherapy in Infants and Toddlers With Low‐Threshold Egg, Milk or Wheat Allergy: A Prospective Cohort Study
por: Katsumasa Kitamura, et al.
Publicado: (2026)
por: Katsumasa Kitamura, et al.
Publicado: (2026)
Antigenicity of proteins in cooked egg powder and skim milk powder for children with egg and milk allergies
por: Michihiro Naito, et al.
Publicado: (2025)
por: Michihiro Naito, et al.
Publicado: (2025)
MEGState: Phoneme Decoding from Magnetoencephalography Signals
por: Suzuki, Shuntaro, et al.
Publicado: (2025)
por: Suzuki, Shuntaro, et al.
Publicado: (2025)
LILAC: Language-Conditioned Object-Centric Optical Flow for Open-Loop Trajectory Generation
por: Kambara, Motonari, et al.
Publicado: (2026)
por: Kambara, Motonari, et al.
Publicado: (2026)
Superprotonic Conduction in Donor Co‐Doped Perovskites
por: Kensei Umeda, et al.
Publicado: (2026)
por: Kensei Umeda, et al.
Publicado: (2026)
Ejemplares similares
-
Open-Vocabulary Mobile Manipulation Based on Double Relaxed Contrastive Learning with Dense Labeling
por: Yashima, Daichi, et al.
Publicado: (2024) -
Mobile Manipulation Instruction Generation from Multiple Images with Automatic Metric Enhancement
por: Katsumata, Kei, et al.
Publicado: (2025) -
ABMAMBA: Multimodal Large Language Model with Aligned Hierarchical Bidirectional Scan for Efficient Video Captioning
por: Yashima, Daichi, et al.
Publicado: (2026) -
Affordance RAG: Hierarchical Multimodal Retrieval with Affordance-Aware Embodied Memory for Mobile Manipulation
por: Korekata, Ryosuke, et al.
Publicado: (2025) -
Task Success Prediction for Open-Vocabulary Manipulation Based on Multi-Level Aligned Representations
por: Goko, Miyu, et al.
Publicado: (2024)