GoCoMA: Hyperbolic Multimodal Representation Fusion for Large Language Model-Generated Code Attribution

Fuente: arXiv
Saved in:
Bibliographic Details
Main Authors: Choudhury, Nitin, Maurya, Bikrant Bikram Pratap, Kuwar, Bhavinkumar Vinodbhai, Buduru, Arun Balaji
Format: Preprint
Published: 2026
Subjects:
Online Access:
Tags: Add Tag
No Tags, Be the first to tag this record!
_version_ 1866910161596055552
author Choudhury, Nitin
Maurya, Bikrant Bikram Pratap
Kuwar, Bhavinkumar Vinodbhai
Buduru, Arun Balaji
author_facet Choudhury, Nitin
Maurya, Bikrant Bikram Pratap
Kuwar, Bhavinkumar Vinodbhai
Buduru, Arun Balaji
contents Large Language Models (LLMs) trained on massive code corpora are now increasingly capable of generating code that is hard to distinguish from human-written code. This raises practical concerns, including security vulnerabilities and licensing ambiguity, and also motivates a forensic question: 'Who (or which LLM) wrote this piece of code?' We present GoCoMA, a multimodal framework that models an extrinsic hierarchy between (i) code stylometry, capturing higher-level structural and stylistic signatures, and (ii) image representations of binary pre-executable artifacts (BPEA), capturing lower-level, execution-oriented byte semantics shaped by compilation and toolchains. GoCoMA projects modality embeddings into a hyperbolic Poincaré ball, fuses them via a geodesic-cosine similarity-based cross-modal attention (GCSA) fusion mechanism, and back-projects the fused representation to Euclidean space for final LLM-source attribution. Experiments on two open-source benchmarks (CoDET-M4 and LLMAuthorBench) show that GoCoMA consistently outperforms unimodal and Euclidean multimodal baselines under identical evaluation protocols.
format Preprint
id arxiv_https___arxiv_org_abs_2604_16377
institution arXiv
publishDate 2026
record_format arxiv
spellingShingle GoCoMA: Hyperbolic Multimodal Representation Fusion for Large Language Model-Generated Code Attribution
Choudhury, Nitin
Maurya, Bikrant Bikram Pratap
Kuwar, Bhavinkumar Vinodbhai
Buduru, Arun Balaji
Computation and Language
Computers and Society
Large Language Models (LLMs) trained on massive code corpora are now increasingly capable of generating code that is hard to distinguish from human-written code. This raises practical concerns, including security vulnerabilities and licensing ambiguity, and also motivates a forensic question: 'Who (or which LLM) wrote this piece of code?' We present GoCoMA, a multimodal framework that models an extrinsic hierarchy between (i) code stylometry, capturing higher-level structural and stylistic signatures, and (ii) image representations of binary pre-executable artifacts (BPEA), capturing lower-level, execution-oriented byte semantics shaped by compilation and toolchains. GoCoMA projects modality embeddings into a hyperbolic Poincaré ball, fuses them via a geodesic-cosine similarity-based cross-modal attention (GCSA) fusion mechanism, and back-projects the fused representation to Euclidean space for final LLM-source attribution. Experiments on two open-source benchmarks (CoDET-M4 and LLMAuthorBench) show that GoCoMA consistently outperforms unimodal and Euclidean multimodal baselines under identical evaluation protocols.
title GoCoMA: Hyperbolic Multimodal Representation Fusion for Large Language Model-Generated Code Attribution
topic Computation and Language
Computers and Society
url https://arxiv.org/abs/2604.16377