Liberal Arts & Reasoning · 6 min read · about 9 min aloud · evidence: Provisional

Technology, AI and Human Judgment

Rules for using AI to sharpen study without surrendering the struggle that builds ability, four tests for any model, and the logic, information theory and critics behind them.

A fluent machine answer feels like knowledge. This page starts with what to do about that: rules for study, and four tests for any model. Then it traces where the ideas came from, and who argued against them.

Using AI as tutor, adversary and tool

The rule: use AI to make your thinking harder, not to excuse it. Good uses:

  • Ask for Socratic questions about a text you have already read.
  • Ask for the strongest objection to your reconstruction of an argument.
  • Ask it to find hidden premises, then check them yourself.
  • Ask for an oral examination, answering aloud before looking anything up.
  • Generate practice problems and compare methods.

Never use it before your first reading of the text. Avoid these bad uses:

  • Replacing the first reading with a summary.
  • Producing notes you have not earned by thinking.
  • Treating a summary as understanding.
  • Deciding anything important without independent verification.

The reason is practical. In Roediger and Karpicke's 2006 experiments, recalling material from memory beat rereading it for long-term retention. That is Established for the memory tasks studied. Extending it to broader skills is a reasonable inference, Provisional. If the tool does the effortful part, you keep the feeling of progress and lose the progress. The routine in The Weekly Loop builds these rules into a weekly rhythm. For coding, see Working with Coding Agents; for tool-using systems, Building Bounded Agents.

Testing a model yourself

Treat any model as a witness to cross-examine. A fluent answer is a fact about the model's training, not evidence that the answer is true. Run these checks:

  1. Hallucination. Ask about obscure but checkable facts. Verify each against a primary source (Ji et al., 2023, survey the causes).
  2. Calibration. Ask for a confidence with each answer. Check whether its 80 percent claims are right about 80 percent of the time. The hook's study (Guo et al., 2017) shows why this matters.
  3. Bias. Swap names, genders or places in an otherwise identical prompt and compare responses.
  4. Prompt sensitivity. Rephrase the same question five ways and note whether the answers change.

Keep a log of results. A handful of your own tests teaches more than any benchmark headline.

How these systems work, roughly

A neural network is a large function with adjustable numbers (weights), tuned to reduce error on examples. Backpropagation (Rumelhart, Hinton and Williams, 1986) computes how to adjust them. Large language models use the transformer (Vaswani et al., 2017). Its attention mechanism lets each word weigh its relation to the others. Models are then often tuned to follow instructions using human feedback (Ouyang et al., 2022). Established as a description of how these systems are built. What happens inside them, and whether it deserves words like "understanding", is Aporetic; honest researchers disagree. "The model learns like a brain" is an analogy: artificial neurons are loose abstractions.

Where the ideas come from

Boole (The Laws of Thought, 1854) showed that reasoning with classes can be written as algebra over two values. That is the seed of digital circuits. Demonstrated (it is mathematics).

Turing's 1936 paper defined a machine that follows simple rules on a tape. He argued it can carry out any procedure we would call mechanical. He also proved that the Entscheidungsproblem, a question posed in logic, has no general algorithmic solution. Demonstrated for that result. The claim that his machine captures all "effective" computation (the Church-Turing thesis) is Established, but it is a thesis, not a theorem.

Shannon's 1948 paper measured information as the uncertainty, or surprise, that a message removes. A fair coin flip carries 1 bit. A coin that lands heads 99 percent of the time carries far less, because the outcome rarely surprises. The entropy formula, H = -sum of p times log2 p, limits compression and reliable communication over a noisy channel. Demonstrated Shannon's "information" measures surprise, not meaning; confusing the two is a classic error.

Wiener (The Human Use of Human Beings, 1950) linked feedback, communication and automation. He worried about machines taking over tasks that used to require human judgment. The formal ideas reappear in the Map of Mathematics and Discrete Mathematics and Graphs.

Limits and critics of machine reason

  • Goedel (1931) proved that any consistent, effectively axiomatised system rich enough for arithmetic contains true statements it cannot prove. Demonstrated Arguments that this shows machines cannot think are Aporetic: the theorem itself says nothing directly about minds. Hofstadter's Goedel, Escher, Bach uses self-reference as an analogy for consciousness; read it as a speculative essay.
  • Dreyfus argued that skilled human action depends on embodied, situational know-how that rule-following programs cannot capture. Today's systems learn from examples instead of following written rules. Whether his argument carries over is Provisional.
  • Weizenbaum saw people confide in his simple ELIZA program. In Computer Power and Human Reason he argued that some tasks should not go to machines, even ones they could do. Those tasks require human responsibility and care.
  • Polanyi observed that much knowledge is tacit: you can ride a bicycle without being able to state the rules. Learning from text alone leaves that residue out.

Other critics aimed at technique itself. Ellul (The Technological Society) argued that modern societies pursue efficiency everywhere until technique justifies itself. Heidegger's essay "The Question Concerning Technology" (1954) claimed technology frames the world as raw material to be ordered and used. Postman (Amusing Ourselves to Death, Technopoly) argued that each medium changes what counts as serious thought. Huxley's Brave New World feared distraction and pleasure; Orwell's Nineteen Eighty-Four feared surveillance and force. Lewis's The Abolition of Man asked whether conquering nature ends in some people controlling others without a shared moral standard. These are interpretive arguments, Provisional at best: questions to test, not conclusions to adopt.

Alignment: the metric is not the good

Systems optimise what they are told to measure. Amodei and colleagues (2016) list the failure modes. They include reward hacking, where a system exploits a loophole in its objective, and negative side effects. The deep point is older than AI. Any proxy can be maximised while the thing it stood for is lost. Test scores, clicks and quarterly targets all fail this way. The incentive problems in Systems, Decisions and Robust Design make the same point. Nutrition research faces this gap too. A flaxseed trial lowered blood pressure, a marker. Its registered count of deaths, strokes and heart attacks ran 5 of 58 against 4 of 52 on placebo. Those are too few events to show a difference. That is why the nutrition evidence ladder ranks biomarkers below trials that count heart attacks and deaths.

What must stay human

Tools shape attention, and attention shapes character. A learner can protect both. Limit notifications and feeds, read long texts away from screens that interrupt, and do some work by hand on purpose. Building a small program or threat-modelling a system gives you evidence about technology that opinion cannot. The choice of ends and the responsibility for outcomes must remain human, the subject of Reasoning for What Ends.

Where this could be wrong

Strongest objection. The study rule leans on experiments in which students recalled or reread prose passages. Those studies did not test AI tutors. A model that asks sharp questions might help more than struggling alone.

Best reply. The rule does not ban AI. It orders the work: attempt first, then ask. That keeps the recall effort the experiments rewarded, and still uses the tool for objections and questions.

What would settle it. A trial comparing draft-first with ask-first learners on the same texts, tested days later, not minutes. Until then the rule stays Provisional.

Try this

  1. Entropy by hand. Compute the entropy in bits of a coin with 90 percent heads. Compare it with a fair coin and explain in a sentence why it is lower.
  2. Cross-examine a model. Run the four tests above with ten questions each. Write one paragraph on where the model was trustworthy and where it was not.
  3. Earn the note. Read a short text and write your own summary from memory. Only then ask an AI to critique it. Record what it caught that you missed.

Further reading

  • Shannon, "A Mathematical Theory of Communication" (1948).
  • Turing, "Computing Machinery and Intelligence" (1950).
  • Russell and Norvig, Artificial Intelligence: A Modern Approach (4th ed., 2020).
  • Weizenbaum, Computer Power and Human Reason (1976).
  • Postman, Technopoly (1992).
  • Lewis, The Abolition of Man (1943).
  • Huyen, AI Engineering (O'Reilly, 2025).

Next

From the SizzlinShred reading shelf. The study page adds a guess-first question, a diagram and 5 check-yourself cards.