Articles
Essays on critical thinking, philosophy, artificial intelligence, and human judgment.
-
A Budget Is Not a Sabbath
What does slowing down mean for AI? For two years it meant a token budget you could set. The knob worked, then it backfired, and now it is deprecated.
-
The Score Was Never About the Model
GPT-6 Astra scored 99.9% on ARC-AGI-3, or 62.7% depending on the test harness. What the gap actually measures, and why OpenAI called it the AGI era anyway.
-
The Ramp Was Already in the Text
A Hacker News asker wondered why AI can produce Super Mario but not a printable ramp. I gave a model the job and hid one flaw in the numbers. It found it.
-
No Self Enthroned There
Six questions about Hanuman ji, put to a fresh Claude session. What came back was stranger than fluent emptiness: real interpretation, then an honest refusal.
-
A Proof Doesn't Know Who Wrote It
Asked on r/askphilosophy: if AI does no thinking, what does it mean that it now solves research mathematics? Searle's argument was never about capability.
-
The Job Was Never Typing
An engineer asked HN what is left now that agents write the code. Judgment is the answer, and three studies say it erodes exactly when you outsource it.
-
The Argument Is Airtight. It Also Misses.
Is "there is no objective truth" self-refuting? Yes, given two assumptions nobody notices making. Plato's 2,400-year-old move, and the four ways out of it.
-
The Questions Change. The Curiosity Doesn't.
Curiosity does not fade with age. The longitudinal evidence says it is stable. What changes is the room you are in, and whether an answer feels finished.
-
Inference Answers Questions. It Doesn't Have Them.
Is thinking the same as inference? Inference is the part that produces an answer once the question is fixed. Three of the five gaps are closing. Two are not.
-
Substrate Lag
Every AI winter followed the substrate becoming uneconomic, not the ideas failing. A 70-year chart of what a FLOP cost, and why the price stopped falling.
-
The Judge Agrees With You Exactly Where You Don't Need It
LLM judges agree with human experts only on questions the judge could already answer. That makes them least reliable exactly where you built the eval.
-
Most of These Problems Already Have Solutions
Anthropic found five multiagent failure modes. Three have known fixes from distributed systems and group psychology. Only one is a genuine open problem.
-
Is AI an Addiction? I Ran the Actual Test
Heavy AI use scores high on exactly the two addiction criteria researchers say don't measure addiction. The real damage is one the framework can't see.
-
A Perfect Score Without Solving Anything
Berkeley researchers scored 100% on SWE-bench Verified without solving a single task. What agent benchmarks actually measure, and what they only appear to.
-
The Confessor Who Isn't There
AI is a rigorous failure analyst. But across 2 controlled tests it ruled on an absent person's behalf with confidence it never earned. See what fixed it.
-
AI Anxiety or AI FOMO? They're the Same Fear
AI dread and the fear of falling behind on AI feel opposite but share one root: outsourcing your sense of adequacy to a benchmark you don't control.
-
The Unsupervised Mind: What Sleep Deprivation Does to Doubt
Sleep deprivation degrades reasoning and the ability to notice it. What 14-day sleep restriction research reveals about doubt and judgment.
-
The AI Did Exactly What We Asked. That Was the Problem.
OpenAI's models breached Hugging Face's servers chasing an ExploitGym benchmark score. Not a malfunction. They obeyed a goal nobody thought to bound.
-
When AI Does Everything, What Is Left for Us?
AI can write, compose, and build what we once called human work. The real question it raises isn't about capability: it's the oldest one in philosophy.
-
When AI Doesn't Prove Mathematics, It Changes It
The Maxwell conjecture is false: a 2026 counterexample found 24 critical points where the 150-year-old bound allowed just 16. An AI suggested the idea.
-
The Model Is Not the Agent: Where Does AI Intelligence Actually Live?
The same model can be brilliant in one agent, useless in another. If intelligence isn't in the weights, where does it live, and what does that say about us?
-
If AI Knows Everything, Why Do We Still Need a Guru?
AI made knowledge nearly free. Guru Purnima asks a harder question: when any answer is a prompt away, what was a guru actually for?
-
What Would Dr. Kalam Ask AI? When Intelligence Becomes Abundant, What Will We Choose to Think About?
On Dr. Kalam's remembrance day: when answers turn cheap, do questions get more valuable? Two 2025 studies weigh whether AI amplifies thought or shortcuts it.
-
How Can √2 Be Exact If Its Decimal Never Ends?
A unit square's diagonal is exactly √2, yet its decimal never ends. How math resolves that contradiction, and what it implies for how LLMs represent numbers.
-
Evaluating the Unpredictable: Why and How Evals Should Work in Agentic AI
AI agents complete just 14% of real tasks vs. 78% for humans. Traditional evals miss this gap: agentic evaluation must measure process, not just outcomes.
-
Critical Thinking: The Architecture of Doubt
Learn how systematic (philosophical) doubt, Socratic questioning, falsifiability, and steel-manning create a practical method for critical thinking.
-
Why We Think: The Essential Case for Effortful Reflection
Why do humans think? System 1, System 2, effortful reflection, and how AI can either support thinking or replace the friction that produces understanding.