<?xml version="1.0" encoding="utf-8" standalone="yes"?><rss version="2.0" xmlns:atom="http://www.w3.org/2005/Atom" xmlns:content="http://purl.org/rss/1.0/modules/content/"><channel><title>Training-Quality on Mohanad Abu-Nayla</title><link>https://hannody.github.io/tags/training-quality/</link><description>Recent content in Training-Quality on Mohanad Abu-Nayla</description><generator>Hugo</generator><language>en-US</language><lastBuildDate>Sat, 30 May 2026 00:00:00 +0000</lastBuildDate><atom:link href="https://hannody.github.io/tags/training-quality/index.xml" rel="self" type="application/rss+xml"/><item><title>Before You Train, Audit: A TinyStories Case Study</title><link>https://hannody.github.io/posts/ai/training/before-you-train-audit-tinystories/</link><pubDate>Sat, 30 May 2026 00:00:00 +0000</pubDate><guid>https://hannody.github.io/posts/ai/training/before-you-train-audit-tinystories/</guid><description>&lt;p&gt;If you are building a small language model from scratch, TinyStories is usually one of the first datasets you reach for. I did too.
Before touching tokenization and training, I ran a quick data audit on the parquet train split. That one step surfaced something worth sharing: a large number of exact duplicate stories sitting quietly in the data.
This post is a practitioner-level writeup, not an accusation. TinyStories is still a solid dataset. The point is just this: verify your data pipeline the same way you verify your model pipeline.&lt;/p&gt;</description></item></channel></rss>