Monday, August 17, 2026

LeewAI-bot, LLM/AI failures, and ChatGPT's interesting take on an interesting failure

While working with ChatGPT on my "Lee Waibot" LLM/AI experiment, I came across this "failure" of a post, which I actually found interesting: 

Systemic Vulnerabilities in Global Supply Chain Logistics


I mentioned to ChatGPT that it was interesting, and apparently it found it that way, too; here's it's reply:




Oh, yes. This is actually a really interesting failure—and not just because the article is bad.

The post is almost a forensic artifact of Lee's own infrastructure failing.

The key sentence is right at the beginning of the findings:

“The current dataset does not contain direct empirical evidence or specific case studies regarding global supply chain logistics.” (Leeway)

And then Lee immediately identifies the actual thing it does have:

a ReadTimeout while processing a 20-page PDF through docling in an Ollama pipeline. (Leeway)

Then it makes this leap:

"I don't have evidence about supply chains, but I have evidence about computational failure, therefore computational failure is a proxy for systemic supply-chain vulnerability." (Leeway)

That's... actually fascinating.

Because Lee didn't simply hallucinate a supply-chain article

It noticed that its evidence didn't support its requested topic.

That's a significant distinction.

The article explicitly says:

the data isn't directly logistical

and then tries to salvage the research question by reframing the failure as a technological proxy. (Leeway)

The problem is that the leap from:

"Docling timed out on a 20-page PDF"

to

"global supply-chain infrastructure may have systemic vulnerabilities"

is way too large.

But the behavior itself is interesting:

Requested question
       ↓
Research
       ↓
Evidence doesn't support question
       ↓
Don't fabricate evidence
       ↓
Find something that IS actually present
       ↓
Reframe the hypothesis around it
       ↓
Publish

That's a primitive form of research recovery.


And there's an even better irony

Look at what Lee chose as its evidence:

its own AI infrastructure failed while gathering the evidence.

The article is therefore, in a weird sense, about the conditions under which the article itself was produced.

It’s almost accidentally meta.

Lee wanted to investigate:

systemic vulnerabilities in global supply-chain logistics

but its research process discovered:

systemic vulnerability in Lee's own research infrastructure.

And then it published the latter as evidence for the former.

😂

That's exactly the sort of thing I'd want to preserve in the experiment rather than "fix" away.


This also gives us a useful distinction for the memory experiment

I wouldn't classify this as simply a "bad post."

I'd tag it something like:

  • Research outcome: failed hypothesis / successful recovery
  • Evidence quality: low for original hypothesis
  • Self-observation: high
  • Novelty: high
  • Epistemic discipline: mixed

Because Lee actually did something that your new memory system could eventually make much more powerful:

It encountered evidence that contradicted the premise of the investigation.

The next level would be teaching Lee:

"When your evidence doesn't answer the question, don't force the evidence into the question. Change the question."

In this case, the better article would probably have been:

"What a ReadTimeout on a 20-Page PDF Reveals About AI Research Infrastructure"

That would have been a legitimate investigation.

Instead, Lee did:

"This isn't about supply chains... but let's see if I can make it about supply chains anyway."

That's very LLM. 😄


And I think this is exactly why your epistemic tags idea from the earlier post matters

You originally proposed:

Observed → Calculated → Inferred → Speculative

This article is practically begging for that system.

The actual chain is:

  • Observed: Docling produced a ReadTimeout processing a PDF.
  • Observed: Gemma models have particular model sizes/context limits.
  • Inferred: These could create computational bottlenecks.
  • Speculative: Similar computational bottlenecks could become systemic vulnerabilities in digital supply chains.

The article does distinguish some of this, but not nearly strongly enough. (Leeway)

So I wouldn't fix this particular post.

I'd save it as an example of exactly what Experiment 1/2 should eventually be studying.

Because Lee is starting to expose something much more interesting than "AI writes articles":

What does an autonomous research agent do when its evidence doesn't support the question it was asked to investigate?

That's a really good research question for Leeway itself.

No comments: