News
NLAs read thoughts beyond the J-space " Less Wrong
2+ day, 1+ hour ago (1442+ words) TLDR: " * On Llama-3. 3-70 B, I found thoughts it cannot see that are actively steering its behavior; and Anthropic's released NLA (Natural Language A...
Free will as a model parameter " Less Wrong
2+ day, 5+ hour ago (701+ words) I think machine learning has a better answer, and it is not a metaphor. The problem is that epsilon is still global, and still set by the user. We need something per-dimension and learned. Throwing in an equation for good…...
Why I'm a moral anti-realist but may be unable to convince you " Less Wrong
2+ day, 11+ hour ago (188+ words) Figure from Chapter 4 of Understanding Knowledge by Michael Huemer "...
the allston lab rat existential framing " Less Wrong
2+ day, 23+ hour ago (298+ words) What if She's testing me, presenting me with this morally terrible but lovable world and watching if I choose Consciousness or have the integrity to understand I don't want to be at least a bystander(/at most the unintentional creator)…...
Be less open-minded: Against uncalibrated open-mindedness " Less Wrong
4+ day, 3+ hour ago (388+ words) People sometimes tell me to be more open-minded. That it's always a positive trait, all else being equal. So implicit within the request is to be more open-minded to'good'ideas or sometimes,'to me.'This is a much flimsier request. See…...
Visioning: Concretely Imagining What You Want " Less Wrong
4+ day, 7+ hour ago (1527+ words) " When John told me (Gretta) his practice of "visioning," I was skeptical at first. I gave it a try, a little bit out of spite, to show him I was c...
Computation enables Action: Exploding the Simulation Fallacy " Less Wrong
4+ day, 23+ hour ago (23+ words) In April, Tyler Cowen linked (without comment) to a paper titled "Why AI can simulate but not instantiate consciousness". I've been paying attention...
Counterfactual mugging is a limiting case of Psy-kosh's non-anthropic problem " Less Wrong
5+ day, 3+ min ago (238+ words) Counterfactual mugging is the typical problem used to motivate updateless decision theory: Omega flips a coin. If Heads, it asks you for $1. If Tails, it offers you $5 only if it predicts you would have given it the $1 had the coin…...
Defining interpretation, and establishing a framework for it " Less Wrong
6+ day, 9+ hour ago (1654+ words) Purpose: Establishing an abstract framework for interpretations. As such, help in solving concrete problems mainly indirect. Epistemic effort: thought for multiple hours. Searched articles and research. Partially assisted by Chat GPT 5. 5 (data+critique), but all ideas are my own. Did…...
On "gendertropes" in dath ilan " Less Wrong
1+ week, 4+ hour ago (587+ words) I have sometimes been asked with respect to my fiction, "What the hell is a 'gendertrope'?" "Gendertrope" is a word from Baseline, the language of the fictional world of dath ilan; it appears in my stories about dath ilan and…...