<?xml version="1.0" encoding="UTF-8"?><rss xmlns:dc="http://purl.org/dc/elements/1.1/" xmlns:content="http://purl.org/rss/1.0/modules/content/" xmlns:atom="http://www.w3.org/2005/Atom" version="2.0" xmlns:itunes="http://www.itunes.com/dtds/podcast-1.0.dtd" xmlns:googleplay="http://www.google.com/schemas/play-podcasts/1.0"><channel><title><![CDATA[Tauseef Ahmad Khan]]></title><description><![CDATA[Nutrition scientist at the University of Toronto. I study sugars, sweeteners and chronic disease, and write about how AI is changing science, learning and the university.]]></description><link>https://blog.tauseefkhan.ai</link><image><url>https://substackcdn.com/image/fetch/$s_!m7T9!,w_256,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F1962d911-9093-469b-832a-8e0be4605a36_835x835.jpeg</url><title>Tauseef Ahmad Khan</title><link>https://blog.tauseefkhan.ai</link></image><generator>Substack</generator><lastBuildDate>Wed, 30 Sep 2026 17:00:02 GMT</lastBuildDate><atom:link href="https://blog.tauseefkhan.ai/feed" rel="self" type="application/rss+xml"/><copyright><![CDATA[Tauseef Ahmad Khan]]></copyright><language><![CDATA[en]]></language><webMaster><![CDATA[tauseefahmadkhan@substack.com]]></webMaster><itunes:owner><itunes:email><![CDATA[tauseefahmadkhan@substack.com]]></itunes:email><itunes:name><![CDATA[Tauseef Ahmad Khan]]></itunes:name></itunes:owner><itunes:author><![CDATA[Tauseef Ahmad Khan]]></itunes:author><googleplay:owner><![CDATA[tauseefahmadkhan@substack.com]]></googleplay:owner><googleplay:email><![CDATA[tauseefahmadkhan@substack.com]]></googleplay:email><googleplay:author><![CDATA[Tauseef Ahmad Khan]]></googleplay:author><itunes:block><![CDATA[Yes]]></itunes:block><item><title><![CDATA[The agentic unslopper: how I stopped my AI grading its own homework]]></title><description><![CDATA[An AI that judges its own writing will usually like it. The fix is a loop that learns from every failure and leaves the final call to a human.]]></description><link>https://blog.tauseefkhan.ai/p/the-agentic-unslopper-how-i-stopped</link><guid isPermaLink="false">https://blog.tauseefkhan.ai/p/the-agentic-unslopper-how-i-stopped</guid><dc:creator><![CDATA[Tauseef Ahmad Khan]]></dc:creator><pubDate>Tue, 29 Sep 2026 18:06:17 GMT</pubDate><enclosure url="https://substackcdn.com/image/fetch/$s_!vFeD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png" length="0" type="image/jpeg"/><content:encoded><![CDATA[<div class="captioned-image-container"><figure><a class="image-link image2 is-viewable-img" target="_blank" href="https://substackcdn.com/image/fetch/$s_!vFeD!,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png" data-component-name="Image2ToDOM"><div class="image2-inset"><picture><source type="image/webp" srcset="https://substackcdn.com/image/fetch/$s_!vFeD!,w_424,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png 424w, https://substackcdn.com/image/fetch/$s_!vFeD!,w_848,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png 848w, https://substackcdn.com/image/fetch/$s_!vFeD!,w_1272,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png 1272w, https://substackcdn.com/image/fetch/$s_!vFeD!,w_1456,c_limit,f_webp,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png 1456w" sizes="100vw"><img src="https://substackcdn.com/image/fetch/$s_!vFeD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png" width="1456" height="816" data-attrs="{&quot;src&quot;:&quot;https://substack-post-media.s3.amazonaws.com/public/images/2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png&quot;,&quot;srcNoWatermark&quot;:null,&quot;fullscreen&quot;:null,&quot;imageSize&quot;:null,&quot;height&quot;:816,&quot;width&quot;:1456,&quot;resizeWidth&quot;:null,&quot;bytes&quot;:164074,&quot;alt&quot;:null,&quot;title&quot;:null,&quot;type&quot;:&quot;image/png&quot;,&quot;href&quot;:null,&quot;belowTheFold&quot;:false,&quot;topImage&quot;:true,&quot;internalRedirect&quot;:&quot;https://tauseefahmadkhan.substack.com/i/218052344?img=https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png&quot;,&quot;isProcessing&quot;:false,&quot;align&quot;:null,&quot;offset&quot;:false}" class="sizing-normal" alt="" srcset="https://substackcdn.com/image/fetch/$s_!vFeD!,w_424,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png 424w, https://substackcdn.com/image/fetch/$s_!vFeD!,w_848,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png 848w, https://substackcdn.com/image/fetch/$s_!vFeD!,w_1272,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png 1272w, https://substackcdn.com/image/fetch/$s_!vFeD!,w_1456,c_limit,f_auto,q_auto:good,fl_progressive:steep/https%3A%2F%2Fsubstack-post-media.s3.amazonaws.com%2Fpublic%2Fimages%2F2ea9c807-4df3-4223-a370-ea3764a46cac_1456x816.png 1456w" sizes="100vw" fetchpriority="high"></picture><div class="image-link-expand"><div class="pencraft pc-display-flex pc-gap-8 pc-reset"><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container restack-image buttonBase-GK1x3M"><svg aria-hidden="true" width="20" height="20" viewBox="0 0 20 20" fill="none" stroke-width="1.5" stroke="var(--color-fg-primary)" stroke-linecap="round" stroke-linejoin="round" xmlns="http://www.w3.org/2000/svg" class="icon-noB79L"><g><path d="M2.53001 7.81595C3.49179 4.73911 6.43281 2.5 9.91173 2.5C13.1684 2.5 15.9537 4.46214 17.0852 7.23684L17.6179 8.67647M17.6179 8.67647L18.5002 4.26471M17.6179 8.67647L13.6473 6.91176M17.4995 12.1841C16.5378 15.2609 13.5967 17.5 10.1178 17.5C6.86118 17.5 4.07589 15.5379 2.94432 12.7632L2.41165 11.3235M2.41165 11.3235L1.5293 15.7353M2.41165 11.3235L6.38224 13.0882"></path></g></svg></button><button tabindex="0" type="button" class="pencraft pc-reset pencraft icon-container view-image buttonBase-GK1x3M"><svg xmlns="http://www.w3.org/2000/svg" width="20" height="20" viewBox="0 0 24 24" fill="none" stroke="currentColor" stroke-width="2" stroke-linecap="round" stroke-linejoin="round" class="lucide lucide-maximize2 lucide-maximize-2 icon-noB79L"><polyline points="15 3 21 3 21 9"></polyline><polyline points="9 21 3 21 3 15"></polyline><line x1="21" x2="14" y1="3" y2="10"></line><line x1="3" x2="10" y1="21" y2="14"></line></svg></button></div></div></div></a></figure></div><p>I wrote AI slop again this morning. An agent drafted, I skimmed, and the text read smoothly enough to pass. Only on a second reading did I see what it was: tidy paragraphs that said less than they seemed to, with the same turns of phrase every few lines. It looked like slop, read like slop and smelled like slop. It belonged in the bin.</p><p>This keeps happening to researchers who use AI seriously, and this week the reason became clearer to me. A team at Meta (Jason Weston, Ilia Kulikov and Swarnadeep Saha) tested whether AI models can tell expert writing from model writing. They gave a model a published paper with one section removed, asked it to write the missing section, and then asked AI models, including the writers themselves, which version was better. The AI judges preferred the model&#8217;s section over the human original in 63.5% to 84.6% of comparisons. When the judges scored against rubrics that a model had written, they preferred the model every time. The work is a <a href="https://facebookresearch.github.io/RAM/blogs/unslop/">blog post</a> and has not yet been peer-reviewed.</p><p>So the most common quality check in agentic writing, asking the AI whether the draft is good, is biased towards yes. The AI is grading its own homework, and it likes its own handwriting.</p><h2>Two kinds of slop</h2><p>Prose slop is the kind of writing everyone notices. In one of my own writeups, an agent&#8217;s draft used the construction &#8220;X, not Y&#8221; nine times. Add a hedge after every claim, lists of three, signposting, and sections that retell the whole paper, and a reader stops trusting the text within a page. Readers can tell it was written by AI, and they switch off. Length is the trigger. Ask a model to improve a sentence or a paragraph and it usually does it very well. Ask it to write a whole essay or paper in one go, and the result is almost always slop.</p><p>The second kind of slop is more dangerous because it looks like care. A draft article of mine said that a study&#8217;s author &#8220;cautions against reading it as proof of harm&#8221;. It sounded balanced. The paper says the opposite: its findings &#8220;provide compelling evidence for the cognitive impacts of AI tool usage&#8221;. A fabricated caveat reads as diligence, so it survives self-review by the AI agent. My own work supplies other examples. In one citation, &#8220;considers&#8221; drifted to &#8220;accounts for&#8221;, turning a design judgment into a statistical adjustment the source never described or intended. A PDF converter turned 0.00247 into 00247. When reading an image, the vision model counted 13 marks on a figure in one run and 19 in the next.</p><p>Asking the model whether its work was good would have caught none of these. Each was caught by checking against something outside the model: the source text, a second extraction engine, a pixel count and my own eyes.</p><h2>A loop that learns</h2><p>I use Claude Code, Codex and several open models to help with my systematic reviews, meta-analyses and papers. Since the summer I have stopped hunting for the perfect prompt and built a loop instead. I call it the agentic unslopper. It has four parts.</p><p>First, every failure becomes a lesson with a tripwire. When I catch an agent&#8217;s mistake, the lesson goes into a shared memory, a folder of short notes that every agent reads, and it must name the check that would have caught it: a search pattern, an audit script or a gate in the workflow. The invented caveat produced such a rule: any sentence saying a source &#8220;cautions&#8221;, &#8220;notes&#8221; or &#8220;warns&#8221; must be found in the source before the draft is saved. The check matters more than the note. Agents forget written advice; a check runs every time.</p><p>Second, numbers never live in the model&#8217;s memory. Every value that is extracted, derived, calculated or estimated is written to disk the moment it exists, and later steps read it from the file. This is a hard rule in the instruction file that all my agents load. Data extraction for a meta-analysis runs through two AI engines plus a visual read of the page: one engine extracts, the second verifies, and a third model re-checks the tables. Any disagreement is resolved against the original PDF. I chose this set after comparing nine extraction engines on my own papers.</p><p>Third, my text is the default. When I edit a draft and an agent reviews my edits, it may flag only what can be verified: a changed number, a certainty claim the study design cannot carry, a citation or a quotation. It may not nudge my wording back towards its own. I learnt this the hard way: I once asked an agent to check and fix the citations in a manuscript, and it also &#8216;fixed&#8217; some of my prose, changing the meaning of what I wanted to convey. I made it a rule this week, after reading the Meta study. A second rule came with it: judge a draft against an expert example of the same kind, and list what is missing instead of giving a score.</p><p>Fourth, the loop feeds itself. Each workflow reads its own lessons at the start of every run, so a mistake made on Monday shapes Tuesday&#8217;s work. This is the recursive part, and the loop itself needs checking. A recent audit found 61 saved lessons about specific workflows, but only 5 in total had been wired into the workflow they concerned. The rest were orphan lessons that no workflow was reading. The agents were learning and then not using what they had learnt. A learning system decays quietly unless you audit the learning as well.</p><h2>What it cost, and what it bought</h2><p>The loop is slower than accepting the first draft. While preparing a perspective on AI in nutrition research for submission, three agents audited every cited claim in parallel, and each source was then checked against the original document. Of 31 cited claims, 24 were fully supported and 7 only partly. We made 13 corrections, most of them to wording that claimed a little more than the source said. None was dramatic; together they were the difference between a paper that is right and one that sounds right.</p><p>Unslopping can also overshoot. When I told an agent that a document I had written with AI&#8217;s help read like AI slop, it cut the text from 2,647 to 1,913 words and removed six of the nine &#8220;X, not Y&#8221; constructions. The result was disjointed and jarring. We restored the transitions and the useful signposts, and settled at 2,037 words. That lesson went into memory too: aim to cut 20&#8211;25%, keep the sentences that carry the argument from one point to the next, and stop there.</p><h2>Where the human stays</h2><p>In July I told my research group that writing is the main choke point in our research projects. AI can write methods and results in our standard laboratory style; the introduction and discussion are where we must put in our own effort. I still think so. A recent experiment by <a href="https://arxiv.org/abs/2607.27191">Kirgis and colleagues</a> reaches the same place from another direction: AI agents completed all the engineering for two research projects without human help, but the projects&#8217; original authors rejected both resulting papers.</p><p>So this is how I now write with AI. I ask the agent for an outline first. I review it, edit it and rewrite the points in my own words, then ask the agent to review it again. When typing is too slow, I switch on the microphone in Claude Code or Codex and talk through what I want to add, and the agent turns it into an outline. Only when I approve the outline do we write prose, one paragraph at a time, and I read and edit every sentence before moving on. I approve each paragraph separately in sequence. In my experience AI handles long-form writing badly: it loses the thread of the argument, and that is where slop creeps in. Reviewing and editing each line oneself is the key. The human mind drives the writing; the AI still remains just a tool that can write.</p><p>The unslopper does not take me out of the loop. My judgment is the gate: the agents stop there and ask for permission, and nothing moves forward until I decide. Because the workflows improve themselves, the output gets better with each round and my decisions get easier. The Meta team trained its model against rubrics built from expert writing. In my system the expert is me, and the rubric is every correction I have made. That is also why I cannot hand all my thinking to the agents. A judge who stops reading goes stale, and then the homework grades itself again.</p>]]></content:encoded></item></channel></rss>