UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification

Document detail

ID

oai:arXiv.org:2410.12318

Topic

Computer Science - Cryptography an... Computer Science - Artificial Inte...

Author

Cai, Jiacheng Yu, Jiahao Shao, Yangguang Wu, Yuhang Xing, Xinyu

Year

2024

listing date

10/23/2024

Keywords

utf fingerprinting model tokens

Metrics

Abstract

Fingerprinting large language models (LLMs) is essential for verifying model ownership, ensuring authenticity, and preventing misuse.

Traditional fingerprinting methods often require significant computational overhead or white-box verification access.

In this paper, we introduce UTF, a novel and efficient approach to fingerprinting LLMs by leveraging under-trained tokens.

Under-trained tokens are tokens that the model has not fully learned during its training phase.

By utilizing these tokens, we perform supervised fine-tuning to embed specific input-output pairs into the model.

This process allows the LLM to produce predetermined outputs when presented with certain inputs, effectively embedding a unique fingerprint.

Our method has minimal overhead and impact on model's performance, and does not require white-box access to target model's ownership identification.

Compared to existing fingerprinting methods, UTF is also more effective and robust to fine-tuning and random guess.

Cai, Jiacheng,Yu, Jiahao,Shao, Yangguang,Wu, Yuhang,Xing, Xinyu, 2024, UTF:Undertrained Tokens as Fingerprints A Novel Approach to LLM Identification

Document

Open

Source

Articles recommended by ES/IODE AI

Computer Science

VR-Splatting: Foveated Radiance Field Rendering via 3D Gaussian Splatting and Neural Points

computer demands rendering neural

The Journal of Neuroscience

Attractor-Like Dynamics Extracted from Human Electrocorticographic Recordings Underlie Computational Principles of Auditory Bistable Perception

low-dimensional ecog subjects neural features dynamics perception

biorxiv

Multiplexed live-cell imaging for drug responses in patient-derived organoid models of cancer

cell organoid patient-derived kinetic system effects cancer pdo models