Full Paper

Synthetic Signal Identification in LLM Datasets

  • Jim Hahn ORCID
  • Penn Libraries, United States
Published
Access and licence
Open access CC BY 4.0
Download PDF

Abstract

This research addresses the quality of training data in LLMs using methods from signaling theory and the talk page metadata of Wikipedia articles. The significance of the method is to lower the cost of information quality assessment in datasets. Natural language processing on metadata text generated sentiment, reading complexity, and self-reference scores as contributions to the computationally derived signals. Results showed that it is possible to understand indicators of information quality using textual computation over the metadata in article pages.

The full text of this article is available as a PDF.

Download PDF

Article details

Published
Section
Full Papers
DOI
10.23106/dcmi.952493684
License
CC BY 4.0 · open access

Described in Dublin Core

This article's metadata, in the vocabulary these proceedings are about.

dcterms:title
Synthetic Signal Identification in LLM Datasets
dcterms:creator
Hahn, Jim
dcterms:date
2024-12-20
dcterms:identifier
doi:10.23106/dcmi.952493684
dcterms:publisher
Dublin Core Metadata Initiative
dcterms:type
Text
dcterms:language
en
dcterms:rights
CC BY 4.0