Search This Blog

Just Where is the Computer that Computes the Universe? (Steven Wolfram's invisible rhetoric)

Do Black Holes warp the universe such that it is self-computable? Kurt Godel famously proved that a computer has to be larger than the problem being computed. This places seemingly fatal constraints on the size of the universe as a computation of itself. Saying as it has become popular to do, that the universe is just one of an infinite set of parallel universe doesn't solve the problem. Even infinities can not be said to be larger than themselves.



TED talk by Stephen Wolfram on the computable universe.

Is it possible that black holes work as Kline bottles for the whole Universe – stretching space-time back around onto itself? If so, it may be possible to circumvent Godel's causal constraints on the computability of the self, as well as the entropic leaking demanded by the second law of thermodynamics. I admit that these questions are not comfortable. They certainly don't result in the kind of ideas I like to entertain. They spawn ideas that seem to be built of need and not logic. They are jokes written to support a punch line.

But something has to give. Either Godel and Turing are wrong, or there is a part of our universe in which they don't apply. There is no other option. If there is a part of the universe not restricted by incompleteness than black holes are obvious candidates if for no other reason than we don't know much about them. I am at once embarrassed by the premise of this thought and excited to talk openly about what is probably the core hiccup in our scientific understanding of the universe. Any other suggestions? At the very least, this problem seems to point to (at least) five options; 1. a deeper understanding of causality will derive Godel and Turning from a deeper causal layer that also has room for super-computable problems. 2. Godel and Turing are dead wrong. 3. the universe is not at all what it seems to be, rendering all of physics mute, and 4. the universe is always in some real way, larger than itself, and 5. evolution IS the computation of the universe, it happens at the only pace allowable by causality, is an intractable program, and can not be altered or reduced, (event cones, the only barriers between parallel simultaneous execution).

I am challenged by the first option, find the second option empirically problematic, am rhetorically repulsed by the third, simply do not know what to do with the fourth, the fifth is where I place my bets but I don't fully understand the implications or the parameters. Personal affinities aside, we had better face the fact that our understanding of the universe is at odds with the universe itself. That we have a set of basic laws that contradict the existence of the universe as a whole is problematic at best. Disturbing.

One of the unknowns that haunt our effort to understand the universe as a system is the ongoing confusion between what we think of as "primary" reality on the one hand and "descriptive" reality on the other. Real or just apparent, it is a distinction that has motivated the clumsily explorations of the "Post-Modern" theoretical movement – it deserves better. I am not so romantic to believe that this dichotomy represents a real qualitative difference between the material and the abstract (made up as it is of the same "real" materials), but this confusion may indeed hint towards a sixth option that, once explained and understood, will obliterate the causal contradictions that have so confused our understanding of the largest of all questions. When a chunk of reality is used as abstraction signifying another part of reality or a part the same reality of which the abstraction is built, does that shift in vantage demand a new physics, a new set of evaluation semantics? What modifications does one have to perform to E = mC^2 when one is computing the physical nature of the equation itself? What new term is to be added to our most basic physical laws such that the causal and the representative can be brought into harmony?

My own view is that the universe, like all systems, like any system, is always in the only configuration it can be in at that time. Wow, that sounds Taoist and I absolutely hate it when attempts at rationality result in assessments that are so easily resonant with emotionally satisfying sentimentality (What the Bleep, and such). But the Second Law clearly points to a maxed out rate as the only possible reading of process at all scales. Computation of anything, including the whole of the universe, is always limping along at the maximum rate dictated by each current configuration. The rate of the process, of the computation, accelerates through time as complexities stack up into self optimized hierarchies of grammar, but the rate is, at each moment, absolutely maxed out.

Are these daft notions chasing silly abstraction-bounded issues or do they point to a real "new [and necessary] kind of science"?

OK, as usual, Mr. Wolfram has expansive dreams – awesomely audacious and attractively resonant notions. Though, from my own perspective, a perspective I will say is more sober and less rhetorical, there are some huge problems that beg to be exposed.

Wolfram's declares: the universe is, at base, computation. Wow, talk about putting the carriage before the horse. That the universe and everything in it is "computing" is hard to dispute. Everywhere there is a difference there will be computation. So long as there is more than one thing, there is a difference. But computation demands stuff. What we call computation is always at base a causal cascade attempting to level an energetic or configurational topology. If you want to call that cascade "computation", well I won't disagree. But no computation can happen unless the running of it diminishes to some extent an energy cline. Computation is slave to the larger more causal activity that is the dissipation of difference. That a universe will result in computation an entirely different assertion.

When Wolfram says that computation exists below the standard model causality that is matter and force, time and space, I am suspicious that he is seeking transcendence, a loophole, access by any means out of the confines of the strictures imposed by physical law. That he is smart and talented and prodigiously effective towards the accomplishment of complex and practical projects does not in itself mean that his musings are not fantastic or monstrous.

Let's play a thought experiment. Let's start from the assumption that Wolfram is correct, that the universe is at base pure computation. His book and this talk hint towards the idea that pure computation running through computational abstraction space, will eventually produce the causality of this universe… and many others. Testing the validity of this assertion is logically impossible. But what we can test is the logical validity of the notion that one could, from the confines of this finite universe, use computation to reach back down to the level of pure computation from which a universe can be made or described. At this level, Turing and Godel both present lock-tight logic showing how Wolfram's assertions are impossible.

In his own examples, Wolfram uses a mountain of human computational space built on billions of years of "computation" (evolution) and technological configurations to make his "simple" programs run. There is NOTHING simple about a program that took a mind like Wolfram's to build (stacked as it is on top of an almost bottomless mountain of causal filtering reaching back to the big bang (or before).

To cover for these logical breaches, Wolfram recites his "computational equivalence" mantra. This is a restating of Alan Turing's notion that a computable problem is computable on any so-called "Turing Complete" computer. But the Turing Machine concept does not contend with the causally important constraint that run-time places on a program. Of course there are non-computable problems. But even within the set of problems that a computer can run and run to completion, there are problems so large that they require billions of times longer to run than the full life cycle of the universe. Problems like these really aren't computable in any practical sense – causality being highly time and location sensitive (isn't that what "causality" means?).

And then there is the parallel processing issue, its potentials and its pitfalls. One might (a universe might), in the course of designing a system that will compute huge programs, decide to break them apart and run sections of the problem on separate machines. Isn't that what nature has done? But there are constraints here as well. Some problems can not be broken apart at all. Some that can, break apart into an unwieldy network constrained by time sensitive links dependent upon fast, wide, and accurate communication channels. if program A needs the result of program B before it can initiate program C but program A only is only relevant for one year and program B takes 2 years to run?

A large percentage of the set of all potential programs, though theoretically run-able on Turing Machines, are not practically run-able given the finite timescales and computational material resource availability. If there is a layer of causality below this universe, and that layer is made of much smaller and much more abundant stuff, than it is conceivable that Godel's strictures on the size of a computer won't conflict with the notion that this Universe could be an example of a Turning Complete computer capable of running the universe as a program.

But Wolfram doesn't stop there. In addition to asserting that a universe is the result of a computation, he says that we humans (and, or, our technology), will be able to write a small program that perfectly computes the universe and that it will be so simple (both as a program and presumably to write it) that we will be able to run it on almost any minimal computer. His cites as example, "rule 30", the fractal equation variation that seems to produce endless variety along an algorithmic theme, as evidence that this universe describing meta-program, is as easy to discover. One has to ask: "Would the running of such a program bud off another universe, or is Wolfram's assertion intentionally restrained to abstraction space?" Given the boldness of his declaration that the universe is a computation, it is reasonable to assume that his statements regarding the discovery of a program that computes a universe is meant in the literal sense. Surely he can talk to the issue of abstraction space vs. causal space, the advantages and constraints of each, and how programs use this difference to compute different types of problems. If he does, he doesn't reveal this understanding to his audience. The distinction between abstraction and causality is slippery and central to the concept of computation.

I am convinced that Stephen Wolfram is so lost in the emotional motivations that push him towards his "computable universe" rhetoric that none of his considerable powers of intellect can save him from the fact that he didn't get the evolution memo. Evolution IS the computation. If it could happen any faster it would have. If he is simply saying that our new understanding of computation will increase the rate and reach of evolution, well then I agree. But if he is saying that our first awkward steps into computation reveal enough of the unknown to expose the God program, the program that will complete all other programs (in a decade), well I can only say that he is nuts.

Stephen is a smart guy. The fact that a mind so capable can overlook, even actively avoid the simple logic that shows terminal flaws in his thesis is yet another reminder of the danger that is hubris. That he never talks to his own motivations, or the potential fallacies upon which his theory depends should be worrisome to anyone listening. I suspect that, like religion, his rhetoric so closely parallels the general human rhetoric, that it will be a rare person who can look behind the curtains and find these logical inconsistencies (no matter how obvious).

I applaud Mr. Wolfram's work. The world is richer as a result. But none of his programming should be taken as guarantee that his theory, at the level of a computational universe is sound.

Randall Reetz

Proactive Fix For Deep Sea Oil Platform Blowouts

If off-shore oil platform developers were required to pre-install a permanent emergency oil blowout collection tent at each wellhead, the disaster unfolding in the gulf of mexico would never have happened.  




The above diagram shows the tent as deployed after a blowout.  Before a blowout the tent would lay flat on the ocean floor in the ready.  When a blowout occurs at the well head (A), a winch (or air filled ballast) (B) pulls the tent up into position over the well head (A).  The tent (C) is composed of an inverted V shaped rigid "tent pole" (D) hinged at pivot points (E) anchored at sea floor.  Once deployed, the tent presents as an inverted pyramid that catches the oil (G) as it rises (oil is lighter than water).  A ten inch hose (H) is lifted from the apex of the tent to the surface of the ocean by buoys (I) along at intervals along its length.  The hose terminates at the surface where a tanker is positioned to pump the oil into its hold until such a time as the well head can be sealed.

Using another approach, the rigid poles are replaced by buoys lifting the apex of the tent.  Four guy lines anchor the tent's corners to the ocean floor.  This option allows for a larger tent and might prove easier to install and deploy.

The entire contraption could also be prebuilt, pre-packaged, and deployed from a GPS guided barge or ship – dropping four anchors or concrete standards at equal radius from the well head and then deploying the collection tent and pumping hose remotely via at-depth gas filled buoys or mechanical winch.

Randall Reetz

Devaluing Survival

The goal of evolution is not survival. Rocks survive much better, longer, and more consistently than biological entities. This should be patently obvious. Survival is a tailing of evolution and achieves a level of false importance probably because those of us doing the observation are so short lived and thus value survival above almost everything else.

In biology as in any other system, evolution is not concerned with nor particularly interested in individual instanciations of a scheme. A being is but a carrier of scheme. And even that is unimportant to THE scheme which can only be one thing – the race towards ever faster and more complete degradation of structure and energy.

To this (or any other) universal end, schemes carry competitive advantage simply and only as a function of their ability to "pay attention to", to abstract, the actual physical grammatical causal structure of the universe. And why is this important? Because a scheme will always have a greater effect on the future of the universe if it "knows" more about the future of the universe. Knowing is a compression exercise. Knowing is two things. 1. acquiring a description of the whole system of which one is a part, and 2. the ability to compress that description to its absolute minimum. A system that does these things better than another system has a greater chance of out-competing its rivals and inserting its "knowledge" into future versions of THE (not "its") scheme. To the extent that an entity pays more attention to its survival (or any other self-centered goal) than to THE scheme, is the extent to which another entity will be able to out-compete it.

Darwin was a great man with an even greater idea (his grandfather Erasmus even more so). But neither had the chops or the context to see evolution at a scope larger than individual living entities or the "species" within which they were grouped competing amongst each other over resources. There was very little understanding of the concept "resources" during his lifetime – certainly not at the meta or generalized level made possible by today's understanding of information and thermodynamics and as a result of Einstein's work its liberation of the symmetry that separated energy, time, distance, and matter. However, Darwin's historically forgivable myopia has out lasted its contextual ignorance and seems instead to be a natural attribute or grand attractor of the human mind. His sophomoric views are repeated ad nauseum to this day.

Randall Reetz

Building Pattern Matching Graphs

I talk a lot about the integral relationship between compression and intelligence.  Here are some simple methods.  We will talk of images but images are not special in any way (just easier to visualize).  Recognizing pattern in an image is easier if you can't see very well.

What?

Blur your eyes and you vastly reduce the information that has to be processed.  Garbage in, brilliance out!



Do this with every image you want to compare.  Make copies and blur them heavily.  Now compress their size down to a very small bitmap (say 10 by 10 pixels) using a pixel averaging algorithm.  Now convert each to grey scale.  Now increase the contrast (about, 150 percent).  Store them thus compressed.  Now compare each image to all of the rest: subtract the target image from the compared image. The result will be the delta between the two. Reduce this combined image to one pixel.  It will have a value somewhere between pure white (0) and pure black (256), representing the gross difference between the two images. Perform this comparison between your target image and all of the images in your data base. Rank and group them from most similar to least.

Now perform image averages of the top 10 percent matches. Build a graph that has all of the source images at the bottom, the next layer is the image averages you just made. Now perform the same comparison to the 10 percent that make up this new layer of averages, that will be your next layer. Repeat until your top layer contains two images. 

Once you have a graph like this, you can quickly find matching images by moving down the graph and making simple binary choices for the next best match. Very fast. If you also take the trouble to optimize your whole salience graph each time you add a new image, your filter should get smarter and smarter.

To increase the fidelity of your intelligence, simply compare individual regions of your image that were most salient in the hierarchical filtering that cascaded down to cause the match. This process can back-propagate up the match hierarchy to help refine salience in the filter graph. Same process works for text or sound or video or topology of any kind. If you have information, this process will find pattern in it. Lots of parameters to tweak. Work the parameters into your fitness or salience breading algorithm and you have a living breathing learning intelligence. Do it right and you shouldn't have to know which category your information originated from (video, sound, text, numbers, binary, etc.). Your system should find those categories automatically.

Remember that intelligence is a lossy compression problem. What to pay attention to, what to ignore. What to save, what to throw away. And finally, how to store your compressed patterns such that the graph that results says something real about the meta-paterns that exist natively in your source set. 

This whole approach has a history of course. Over the history of human scientific and practical thought many people have settled in on the idea that fast filtering is most efficient when it is initiated on a highly compressed pattern range. It is more efficient for instance to go right to the "J's" than to compare the word "joy" to every word in a dictionary or database. This efficiency is only available if your match set is highly structured (in this example, alphabetically ordered). One can do way way way better than alphabetically ordered lists of 3 million words. Lets say there are a million words in a dictionary. If one sets up a graph, an inverted pyramid, where each level where the level one has 2 "folders" and each folder is named for the last word in the subset of all words at that level divided into two groups. The first folder would reference all words from "A" to something like "Monolith" (and is named "Monolith") The second folder at that level contains all words alphabetically larger than "Monolith" (maybe starting with "Monolithic") and is named "Zyzer" (or what ever the last word is in the dictionary). Now, put two folders in each of these folders to make up the second tier of your sorting graph. At the second level you will have 4 folders. Do this again at the third level and you will have 8 folders each named for the last word in the graph referenced in the tiers of the graph above them. It will only take 20 levels to reference a million words, 24 levels for 15 million words. That represents a 6 order of magnitude savings over an unstructured sort. 

A cleaver administrative assistant working for Edward Hubble (or was it Wilson, I can't find the reference?) made punch cards of star positions from observational photo plates of the heavens and was able to perform fast searches for quickly moving stars by running knitting needles into the punch holes in a stack of cards.



Pens A and B found their way through all cards. Pen C hits the second card.

What matters, what is salient, is always that which is proximal in the correct context. What matters is what is near the object of focus at some specific point in time.

Lets go back to the image search I introduced earlier. As in the alphabetical word search just mentioned, what should matter isn't the search method (that is just a perk), but rather the association graph that is produced over the course of many searches. This structured graph represents a meta-pattern inherent in the source data set. If the source data is structurally non-random, its structure will encode part of its semantic content.  If this is the case, the data can be assumed to have been encoded according to a set of structural rules themselves encoding a grammar.

For each of these grammatical rule sets (chunking/combinatorial schemes) one should be able to represent content as a meta-pattern graph. One of the graphs representing a set of words might be pointers to the full lexicon graph. A second graph of the same source text might represent the ordered proximity of each word to its neighbors (remember the alphabetical meta-pattern graph simply represents the neighbors at the character chunk level).

What gets interesting of course are the meta-graphs that can be produced when these structured graphs are cross compressed. In human cognition these meta-graphs are called associative memory (experience) and are why we can quickly reference a memory when we see a color or our nose picks up a scent.

At base, all of these storage and processing tricks depend on two things, storing data structures that allow fast matching, and getting rid of details that don't matter. In concert these two goals result in a self optimization towards maximal compression.

The map MUST be smaller than the territory or it isn't of any value.

It MUST hold ONLY those aspects of the territory that matter to the entity referencing them. The difference between photos and text: A photo-sensor in a digital camera doesn't know for human salience. It sees all points of the visual plane as equal. The memory chips upon which these color points are stored see all pixels as equal. So far, no compression, and no salience. Salience only appears at the level of where digital photos originate (who took them, where, and when). On the other hand, text is usually highly compressed from the very beginning. What a person writes about and how they write it always represents a very very very small subset of 

Compression as Intelligence (Garbage Out, Brilliance In)

I am convinced that the secret to developing intelligence (in any substrate, including your brain) lies in the percentage of the data coming in that you are willing (or forced) to toss. Lossy compression is the key to intelligence. Of course there is a caveat… you can't just trash anything and everything.

The first line of the book I am writing about evolution: "What matters is what matters, knowing what matters and how to know it matters the most."

I am convinced that evolving systems can only work towards mechanisms that process salience if they are forced to maximize the amount of stuff they can trash.

If you are forced to get rid of 99.999 percent of everything that comes in, well you will have to get good at knowing the difference between needles and hay and you will have to get good at knowing the difference in a hurry. The "needles and hay" metaphor doesn't map well to what I am talking towards. If the system you are dealing with is so unstructured as to fit the haystack metaphor, you really aren't doing anything I would classify as intelligence. If there is nothing of structure in the haystack you are storing than your compression system should already have tossed the whole thing out.

Many techniques for the filtering of essence, for finding pattern, for storing pattern and for storing pattern of pattern have been developed. The most impressive reduce raw input streams and store pattern from the most general to the most specific as hierarchically stratified graphs.

Being forced to reduce data to storage formats that maximize lossy-ness minimizes necessary storage. But that is just a perk. What really gates intelligence is the amount of a complex system (or map thereof) that can be made proximal to immediate processing. Our brains might be big and mighty, but what really matters is how much of the right parts of what is stored can be brought together in one small space for semi-real-time simulations processing. Information, when organized optimally for maximal storage density, will also be information that is ideally organized for localized serialization and simultaneity of processing.

To think, a system has to be able to grab highly compressed pattern hierarchies and move them into superposition on top of each other for near instantaneous comparison. You can't do this with a whole brain's worth of data, no matter how well organized it is.

Lets say you have to store everything you know about every sport you have ever heard of, and you have to do it in a very limited space. You will be forced to build a hierarchy of grammars in which general concepts shared in every sport (opponents, the goal to win, a set of rules and consequences, physical playing geometries, equipment, etc.), with layers of groupings that allow for the similarities between some sports and so on up to the specifics that are are only present in each individual sport. Keep compressing this set. Always compress. Try all day (or all night) for even more compression. Compress until you can't even get to lots of the specifics any more. Keep compressing. Dump the sports you don't care about. Keep on throwing stuff out.

Now lets say I have some sort of morbid sense of humor and I tell you that you are going to have to store everything you encounter and everything you think about, your entire life, in that same database that you have optimized for sports.

You will have to learn to look for the meta-patterns that will allow you to store your first romance in a structure that also allows you to store everything you know about kitchen utensils and geo-politics and the way the Beatles White Album makes you feel when it is windy outside.

The necessity to toss, enforced by limited storage and an obsession to compress will result in domain-blending salience hierarchies. It is why we can find deep similarities between music and geological topologies. It is why we can "think".

For years people have tried to come up with the algorithms of thought. What we need instead is to build into our artificial systems, a very mean and ornery compression task master that forces over time, all of our disparate sensation streams into the same shared graph.

Once you have all of your memories stored within the same graph, by necessity sharing the same meta-pattern, the job of evolving processing algorithms is made that much easier.

An intelligent system will spend most if not all of its time compressing data. We have a tendency to bifurcate the behavior of a mind into storage on the one hand, and processing on the other. I am beginning to think that the thing we call "thinking" and "thought" is exclusively and only a side-effect of constant attempts at compression – that there really isn't anything separate that happens outside of compression. Is this possible?

Randall Reetz