Abstract
This research addresses the quality of training data in LLMs using methods from signaling theory and the talk page metadata of Wikipedia articles. The significance of the method is to lower the cost of information quality assessment in datasets. Natural language processing on metadata text generated sentiment, reading complexity, and self-reference scores as contributions to the computationally derived signals. Results showed that it is possible to understand indicators of information quality using textual computation over the metadata in article pages.
The full text of this article is available as a PDF.
Download PDFArticle details
- Published
- Section
- Full Papers
- Published in
- DCMI-2024 Toronto, Canada Proceedings
- License
- CC BY 4.0 · open access
- Download
- Download PDF
Described in Dublin Core
This article's metadata, in the vocabulary these proceedings are about.
- dcterms:title
- Synthetic Signal Identification in LLM Datasets
- dcterms:creator
- Hahn, Jim
- dcterms:date
- 2024-12-20
- dcterms:identifier
- doi:10.23106/dcmi.952493684
- dcterms:isPartOf
- DCMI-2024 Toronto, Canada Proceedings
- dcterms:publisher
- Dublin Core Metadata Initiative
- dcterms:type
- Text
- dcterms:language
- en
- dcterms:rights
- CC BY 4.0