03
A Photo Costs Energy

In 1937, a twenty-year-old Claude Shannon defended a master’s thesis later called “possibly the most important master’s thesis in history.” He proved something engineers shrugged at first — and then built digital civilization on. The gist: an electrical switch is a “yes” or a “no.” Two switches give “yes if both are yes” or “yes if at least one is yes.” From these combinations of “yes” and “no” you can build any logical operation. The entire world of electronics — from a light bulb to a smartphone — is logic written in wires.
Shannon had no idea he’d just laid the foundation of digital civilization. He was just solving a problem for the phone company.
Eleven years later he published a paper that created information theory — an entire science, invented by one person in one summer. Shannon introduced the concept of a “bit” — the minimum unit of information, the answer to one yes-or-no question. And he showed that everything that can be transmitted — from voice to video, from genetic code to thought — is described by a single formula. (The formula: . You don’t have to understand it. You just have to know it exists.)
His boss at Bell Labs reportedly said: “This is brilliant, but who needs it?” Who needed it became clear fifty years later: everyone. Every byte, every pixel, every stream — Shannon. But what’s ahead is a story of an entirely different genre, in which the bit goes from being a unit of data to being a unit of physics itself, and the people surprised by it are no longer telephone engineers but Nobel laureates.

In 1961, IBM physicist Rolf Landauer proved something strange: erasing information produces heat.
It produces it, not “can produce” — inevitably, by a law of physics. Every time you erase one bit — one “yes” or “no” — the Universe pays in energy; a negligible amount (roughly a trillion trillion times less than it takes to lift a speck of dust one millimeter), but strictly greater than zero and always.
For fifty-one years this result was considered a theoretical curiosity — elegant but useless, the way proving an ideal gas doesn’t exist is useless: correct, but who cares?
In 2012 Antoine Bérut’s group (Nature) experimentally confirmed Landauer’s principle, measuring the heat from erasing a single bit — and it matched the prediction exactly.
And here’s why it matters: if erasing information produces a physical effect, then information is physical directly and literally; it is physics. A bit is just as much a physical quantity as a joule or a meter.
The consequence you want to wave away but can’t: every time you delete your ex’s photo from your phone, the Universe pays an energy cost. A tiny one — joules per bit — but a real one. Information doesn’t disappear for free — not anywhere, not ever, not even on your phone.
John Archibald Wheeler arrived at this conclusion earlier — and from a completely different direction.
Wheeler was a central player in twentieth-century physics. He developed the model of nuclear fission with Bohr. He introduced the term “black hole.” His grad student was Richard Feynman. When Wheeler says something, physics listens — even when it doesn’t want to, because ignoring him in this profession costs considerably more than hearing him out.
In 1990, near the end of his career, Wheeler formulated a program of three English words — “It from bit” — which translates as “Everything from information” and sounds like a slogan but is in fact a technical hypothesis.
The full formulation: “Every it — every particle, every field, even spacetime itself — derives its function, its meaning, its very existence entirely — though in some contexts indirectly — from the answers to yes-or-no questions, from binary choices, from bits.”
Wheeler meant something far more radical than a computer metaphor: information is more fundamental than matter, and an atom is “information manifesting as a thing,” rather than the other way around.
The difference here is decisive. In the first formulation, information is a tool of description, like a ruler or a scale; in the second, it becomes a substance — the stuff the world is made of.

Wheeler got there through quantum mechanics and, to be precise, through the delayed-choice experiment.
The standard double-slit experiment: a photon passes through two slits. If we don’t observe which slit it went through, the screen shows an interference pattern — the photon “went through both.” If we do observe, the interference vanishes — the photon “went through one.” This is standard quantum mechanics, known since the 1920s.
Wheeler proposed a twist: the decision to observe or not is made after the photon has already passed through the slits. By classical logic the decision came too late — the photon had already “chosen” which slit to go through. But quantum mechanics says: no. The observer’s decision now determines which slit the photon went through then.
This isn’t theory. The experiment was realized by Jacques et al. (Science, 2007). The result matched Wheeler’s prediction. A decision in the present affects an event in the past.
Stop and reread that: the decision — now, the event — then, the cause — after the effect. This is an experimental fact, published in Science and replicated.
What does this mean for “time”? That time might be a computational order — a sequence of resolutions in a graph, which, under certain conditions, can run in both directions.
If reality is an information graph (and we’ll arrive at this in chapter eleven), then no link in the graph has any built-in direction of time: it has a source and a target, and the source can perfectly well be “later” than the target. Time in such a picture is emergent, arising from the statistics of computations much the way temperature arises from molecular motion.
Wheeler understood this intuitively; Vanchurin formalized it. Our “sense” of linear time is, possibly, the result of a filter that orders computations into a sequence for the convenience of the “user” (us); under the hood, the order isn’t obliged to be linear.
This could explain déjà vu, premonitions, and “folds in reality”: an event that “hasn’t happened yet” has already been resolved at a deep level, it’s just that its resolution hasn’t yet rendered on the surface.
Non-locality of the information graph — that’s the whole explanation.
Wheeler called this the “participatory universe”: the observer doesn’t passively record reality but participates in creating it. Every act of observation is an answer to a yes-or-no question, and that answer is a bit. And from these bits, reality is assembled.
It from bit.
Anton Zeilinger — the 2022 Nobel laureate — developed this idea into a concrete principle. His proposal sounds simple: the smallest quantum system — a single “qubit” — carries exactly one bit of information, exactly one “yes or no,” no more and no less. And from this constraint, Zeilinger claims, follow all the weirdnesses of quantum mechanics — literally all of them, without exceptions.
Why can’t you simultaneously know where a particle is and where it’s going? Because one bit can’t answer two questions at once. Like a coin: it can land heads or tails, but not both at once.
Why does a particle’s state “collapse” into one specific outcome when you measure it? Because you asked a question — and the bit answered. Before the question, the answer did not exist. At all. Anywhere. Like a page in a book that gets written the moment you open it.
Why do two “entangled” particles on opposite ends of the Universe instantly “know” about each other? Because they are not two. They are one bit of information stretched across two locations. Not two objects connected by an invisible string. One object with two addresses.
Zeilinger didn’t prove “it from bit” in the strict sense, but he did something almost equivalent: he showed that once you accept information as the foundation, the whole of quantum mechanics suddenly stops being “weird” and turns into the only logically tidy physics there is — the one that works completely only in a world whose substance turns out to be informational by nature.

With holography, the thesis “information is primary” finally stops being a philosophical declaration and acquires blueprints: what Wheeler was assembling out of quantum bits and intuition, Bekenstein, ’t Hooft, and Maldacena wrote down in surface areas, projections, and equations — and from that moment information acquired, for the first time, its own geometry, and a unit of measurement to go with it.
In 1973 the physicist Jacob Bekenstein derived a formula that left his colleagues walking around with their mouths open for a week. A black hole — a region from which not even light can escape — stores information. How much? Bekenstein did the math and got an answer that wouldn’t fit in the head: the amount of information stored is determined by the surface area of the black hole, not by its inner volume — precisely area, not volume, as common sense would demand.
It’s like finding out that the contents of your apartment — furniture, books, the cat — are determined by the wallpaper, and not “reflected” in it but precisely determined. Know the wallpaper — you know everything inside.
Gerard ’t Hooft (1993) and Leonard Susskind (1995) generalized this to the entire Universe. The holographic principle: all the information contained in any volume of space can be fully described by a theory on its boundary — meaning the three-dimensional world turns out to be a projection of two-dimensional information, a hologram in the most literal physical sense of the word.
Juan Maldacena proved this rigorously in 1997. The details are for specialists (and they fill two hundred pages). The gist is for everyone: Maldacena showed that the physics inside a certain space is fully and exactly described by a theory on its boundary. The volumetric world turns out to be a projection of a surface in the strictest mathematical sense the word can carry — and Maldacena’s paper remains the most cited in theoretical physics to this day (over twenty-five thousand citations on INSPIRE-HEP as of 2025), and physics hands out citation counts like that only to theorems, and a theorem is what Maldacena proved.
What does this mean for us, readers who have never dived into a single black hole?
It means that “depth” is an illusion: volume, distance, and three-dimensional space are computed from information on a surface, and spacetime itself precipitates out as a finished projection — what ordinary physics took for raw material turns out, in the holographic language, to be merely the result of processing.
Here’s an analogy that will make this concrete.
In 2020 DeepMind taught the neural network AlphaFold to solve a problem biologists had struggled with for fifty years: given a gene’s record — a one-dimensional string of letters — predict the three-dimensional shape of a protein. Input: text. Output: object. Between them: rules by which the text folds into form. AlphaFold doesn’t “build” the protein — it “folds” information into matter.
Now imagine “AlphaFold for reality”: a process that folds streams of information into stable forms. Input — code. Folding rules — the laws of physics. Output — matter, space, time. In this picture, the laws of nature turn out to be assembly instructions, thanks to which the informational chaos has time to fold into a stable world before it shakes itself back into random flickers; there are no external tablets in the system on which someone wrote those laws down, because a self-computing reality simply has no “outside” from which any such tablets could have come.
In this model, the laws of nature stop being an external given and become protocols of correct folding — the instructions by which informational chaos folds into a stable fabric of the world and holds itself in that stability appreciably longer than a single iteration.
In empty space — even in absolute vacuum — tiny disturbances constantly “flicker”: particles that are born from nothing and immediately vanish. Physicists call these quantum fluctuations. In our analogy, these are assembly errors. Information tries to fold into a stable form, but not every attempt succeeds: short-lived forms appear that immediately fall apart.
Across all scales, one principle operates: a system can “fold wrong.” If a wrong configuration turns out to be more stable than the right one, it locks in. In biology this gives us prions: proteins that folded incorrectly and now force other proteins to fold the same way (this causes Alzheimer’s, mad cow disease). In thinking, obsessive thoughts: a pattern that “got stuck” and loops endlessly. In cosmology, disturbances in the structure of space that, once frozen, became... galaxies.
The fabric of the Universe holds together through a delicate balance between correct folds and errors, and, possibly, it’s the errors that turn out to be the real engine of evolution: an error-free world would be a perfect copying machine endlessly reproducing one and the same configuration — a world in which there would simply be nowhere for the new to come from.
Before going further, an honest book about information is obliged to say one unpleasant thing — the one without which even the most beautiful chapter on the computational nature of the world inevitably slides into a slogan with a physical accent.
If the world is a computation, then computation has limits, and those limits have long been proven in the same dry mathematical sense in which the Pythagorean theorem is proven — so arguing with them makes about as much sense as arguing with the fact that a triangle’s angles add up to a hundred and eighty degrees.
Gödel, 1931. Kurt Gödel was twenty-five when he proved something mathematics still hasn’t recovered from. The gist: in any sufficiently complex system of rules (and “sufficiently complex” already means arithmetic — addition and multiplication) — there inevitably exist statements that are true but that are impossible to prove from inside the system. There is truth that the rules can’t reach, no matter how good the rules are.
For mathematics this was a blow. The great Hilbert, starting with his Paris Address of 1900 (the second of the famous 23 Problems — the consistency of arithmetic) and developed at length through the 1920s, set the goal of formalizing all of mathematics and proving it was consistent. Gödel showed that this goal is mathematically unreachable forever: inside arithmetic itself, there is simply no route by which one could get to its own consistency, and no matter how much diligence future mathematicians pour into the search, they won’t find such a route, because it isn’t there.
Turing, 1936. Alan Turing was twenty-four when he asked a simple question: “Can you write a program that, given any other program, tells you whether it will hang or not?” The answer: you can’t. Not on any computer. Not now, not in a million years, not on a computer the size of the Universe. The difference between this Turingian “you can’t” and the optimistic “we just haven’t figured it out yet” is the same as the difference between “haven’t yet found the key to the door” and “the door in this wall never existed; the wallpaper there is just painted to look like a door, and even under a microscope there isn’t a keyhole.”
Chaitin, 1975. Gregory Chaitin defined the number — the probability that a randomly chosen program will eventually finish. The number exists, it is definite and finite, but at the same time it is uncomputable: every digit after the decimal point is a mathematical truth that no rules can reach, and itself turns out to be built-in randomness in the strongest sense of the word — knowledge in that place simply ends, the way a map ends at the edge of the world, beyond which there isn’t “unexplored territory” anymore: there’s simply nothing.
Why am I telling you all this — Gödel, Turing, and Chaitin — here?
Because if the Universe is a computational graph (and we’re heading there), it inherits these limitations. Inevitably. Which means:
First. From inside the system you cannot prove all the truth about the system, and we are inside, from which it follows with mathematical directness that some truths about reality are inaccessible to us in principle: an eye unable to see its own pupil — a metaphor that in our case turns out to be a strict theorem.
Second. The system inevitably contains processes for which it’s impossible to predict whether they will terminate — and this is a structural property of any sufficiently rich computational system, with no relation whatsoever to the layman’s “chaos” or to the philosopher’s “free will.”
Third. There exists fundamental randomness — a structural property of the computational substrate, in principle irreducible to “lack of knowledge”; quantum uncertainty may be a manifestation of this randomness — Chaitinian noise built into the architecture of the world, which stops looking like a mystery the moment you stop trying to explain it.
The conclusion is sobering: even if it turns out that information is the foundation of reality, every bit is physical per Landauer, and the Universe really is a self-computing graph, still a complete description of that graph from inside the graph itself does not exist, and this is a mathematically proven property of reality itself, to which future physics will be unable to add anything, no matter how ambitious its graduate students.
Any “theory of everything” will necessarily be incomplete — a claim that only at first glance sounds pessimistic, and on second turns out to be a theorem checkable by schoolbook methods over one long evening.
This book knows it is incomplete — for the simple reason that a complete book about reality is mathematically impossible, and it’s with this calm acknowledgment that any honest map begins, before the first line appears on it.
Vlatko Vedral, an Oxford professor, wrote a book with a telling title — “Decoding Reality: The Universe as Quantum Information” (2010) — and his argument runs as follows: information is a more fundamental category than matter or energy, because both matter and energy can be fully translated into the language of information, whereas the reverse translation — information into the language of matter or energy — inevitably loses something essential.
Seth Lloyd (MIT) did the math: the Universe has performed at most computational operations over its lifetime and stores bits (Physical Review Letters, 2002). The Universe is — literally — computing itself.
Computing literally: every physical interaction is an exchange of information, every quantum process is a computation. Physics turns out to be information processing on a specific substrate, and the question “what is the substrate made of?” may be the wrong one to ask, because the substrate is also information. You get a turtle standing on another turtle, standing on another...
Wheeler would say: “The turtles end. At the bottom — bits.”
Why does any of this matter for our story — for understanding why the brain and the Universe look the same?
Here’s why: if information is fundamental and matter is derived, the structural similarity between the brain and the Universe stops being a coincidence. Both systems turn out to be manifestations of the same informational principles, lying deeper than physics and biology — because physics and biology themselves are manifestations of information.
There is something deeper than physics and deeper than biology that determines the architecture of both — and that something is information: structured, integrated, self-organizing.
In the next chapter we’ll meet the person who went further than anyone and said it straight: “The Universe is a neural network” — literally, mathematically, with quantum mechanics and gravity derived from a single formalism.
His name is Vitaly Vanchurin, and physics is still digesting his 2020 paper — or rather not so much digesting as slowly deciding which drawer to put it in.