Let me take a stab at defending compression as equivalent to intelligence.
Standard string compression (LZW, etc.) works by understanding and then exploiting the sequencing rules that result in the redundancy built into most (all?) languages and communication protocols.
Compression is necessary in any storage/retrieval/manipulation system for the simple reason that all systems are finite. Any library, any hard drive, any computer memory… all finite. If working with primary in-situ environments was as efficient as working with maps or abstractions we would never have to go through the trouble of making maps or abstracting and filtering and representing.
It might seem sarcastic even to say it, but a universe is larger than a brain.
You have however stumbled upon an interesting insight. Where exactly is intelligence? In classic Shannon information theory, and the communication metrics (signal/noise ratio) upon which it is based, information is a duality where data and cypher are interlocked. In this model, you can reduce the size of your content, but only if you increase the size (or capacity) of the cypher. Want to reduce the complexity of the cypher, well you are forced to accept the fact that your content will grow in size or complexity. No free lunch!
In order to build a more robust cypher, one has to generalize in order find salience (the difference that make a difference) in a greater and greater chunk of the universe. It is one thing to build an data crawler for a single content protocol, quite another to build a domain and protocol independent data crawler. It is one thing to build hash trees based on word or token frequency and quite another to build them based on causal semantics (not how the words are sequenced, but how the concepts they refer to are graphed.
I think the main trouble you are having with this compression = intelligence concept has to do with a limited mapping of the word "compression".
Lets say you are driving and need to know which way to turn as you approach a fork in the road. If you are equipped with some sort of mental abstraction of the territory ahead, or on a map, you can choose based on the information encoded into these representations. But what if you didn't? What if you could not build a map, either on paper, or in your head. Then you would be forced to drive up each fork in turn. In fact, had you no abstraction device, you would have to do this continually as you would not be able to remember the first road by the time you took the second.
What if you had to traverse every road in every city you came to just to decide which road you were meant to take in the first place? What if the universe it self was the best map you could ever build of the universe? Surely you can see that a map is a form of compression.
But lets say that your brain can never be big enough to build a perfect map of every part of the universe important to you. Lets imagine that the map-building map you build in order to create mental memories of roads and cities is ineffective at building maps of biological knowledge or physics or the names and faces of your friends. You will have to go about building unique map builders for each domain of knowledge important to you. Eventually, every cubic centimeter of your brain will be full of domain-specific map making algorithms. No room for the maps!
What you need to build sited is a universal map builder. A map builder that works just as well for topological territory as it does for concepts and lists and complex n-dimensional pattern-scapes.
Do so and you will end up with the ultimate compression algorithm!
But your point about where the intelligence lies is important. I haven't read the rules for the contest you sight, but if I were to design such a contest, I would insist that the final byte count of each entrants' data also include the byte count of the code necessary to unpack it.
I realize that even this doesn't go far enough. You are correctly asserting that most of the intelligence is in the human minds that build these compression algorithms in the first place.
How would you go about designing a contest that correctly or more accurately measures the full complexity of both cypher and the content it interprets?
But before you do, you should take the time to realize that a compression algorithm becomes a smaller and smaller component of the total complexity metric the more often it is used. How many trillions of trillions of bytes have been trimmed from the global data tree over the lifespan of use of MPEG or JPEG on video and images? Even if you factor in a robust calculation of the quantum wave space inhabited by the humans brains that created these protocols it is plain to see that use continues to diminish the complexity contribution of the cypher no matter how complex.
Now what do you think?
Randall Lee Reetz
Change increases entropy. The only variable; how fast the Universe falls towards chaos. Determining this rate is the complexity being carried. Complexity exists only to increase disorder. Evolution is the refinement of a fitness metric. It is the process of refining a criteria for the measurement of the capacity of a system to maximize its future potential to hold complexity. This metric becomes ever more sophisticated, and can never be predetermined. Evolution is the computation.
Search This Blog
Showing posts with label intelligence. Show all posts
Showing posts with label intelligence. Show all posts
Building Pattern Matching Graphs
I talk a lot about the integral relationship between compression and intelligence. Here are some simple methods. We will talk of images but images are not special in any way (just easier to visualize). Recognizing pattern in an image is easier if you can't see very well.
What?
Blur your eyes and you vastly reduce the information that has to be processed. Garbage in, brilliance out!
Do this with every image you want to compare. Make copies and blur them heavily. Now compress their size down to a very small bitmap (say 10 by 10 pixels) using a pixel averaging algorithm. Now convert each to grey scale. Now increase the contrast (about, 150 percent). Store them thus compressed. Now compare each image to all of the rest: subtract the target image from the compared image. The result will be the delta between the two. Reduce this combined image to one pixel. It will have a value somewhere between pure white (0) and pure black (256), representing the gross difference between the two images. Perform this comparison between your target image and all of the images in your data base. Rank and group them from most similar to least.
Now perform image averages of the top 10 percent matches. Build a graph that has all of the source images at the bottom, the next layer is the image averages you just made. Now perform the same comparison to the 10 percent that make up this new layer of averages, that will be your next layer. Repeat until your top layer contains two images.
Once you have a graph like this, you can quickly find matching images by moving down the graph and making simple binary choices for the next best match. Very fast. If you also take the trouble to optimize your whole salience graph each time you add a new image, your filter should get smarter and smarter.
To increase the fidelity of your intelligence, simply compare individual regions of your image that were most salient in the hierarchical filtering that cascaded down to cause the match. This process can back-propagate up the match hierarchy to help refine salience in the filter graph. Same process works for text or sound or video or topology of any kind. If you have information, this process will find pattern in it. Lots of parameters to tweak. Work the parameters into your fitness or salience breading algorithm and you have a living breathing learning intelligence. Do it right and you shouldn't have to know which category your information originated from (video, sound, text, numbers, binary, etc.). Your system should find those categories automatically.
Remember that intelligence is a lossy compression problem. What to pay attention to, what to ignore. What to save, what to throw away. And finally, how to store your compressed patterns such that the graph that results says something real about the meta-paterns that exist natively in your source set.
This whole approach has a history of course. Over the history of human scientific and practical thought many people have settled in on the idea that fast filtering is most efficient when it is initiated on a highly compressed pattern range. It is more efficient for instance to go right to the "J's" than to compare the word "joy" to every word in a dictionary or database. This efficiency is only available if your match set is highly structured (in this example, alphabetically ordered). One can do way way way better than alphabetically ordered lists of 3 million words. Lets say there are a million words in a dictionary. If one sets up a graph, an inverted pyramid, where each level where the level one has 2 "folders" and each folder is named for the last word in the subset of all words at that level divided into two groups. The first folder would reference all words from "A" to something like "Monolith" (and is named "Monolith") The second folder at that level contains all words alphabetically larger than "Monolith" (maybe starting with "Monolithic") and is named "Zyzer" (or what ever the last word is in the dictionary). Now, put two folders in each of these folders to make up the second tier of your sorting graph. At the second level you will have 4 folders. Do this again at the third level and you will have 8 folders each named for the last word in the graph referenced in the tiers of the graph above them. It will only take 20 levels to reference a million words, 24 levels for 15 million words. That represents a 6 order of magnitude savings over an unstructured sort.
A cleaver administrative assistant working for Edward Hubble (or was it Wilson, I can't find the reference?) made punch cards of star positions from observational photo plates of the heavens and was able to perform fast searches for quickly moving stars by running knitting needles into the punch holes in a stack of cards.
Pens A and B found their way through all cards. Pen C hits the second card.
What matters, what is salient, is always that which is proximal in the correct context. What matters is what is near the object of focus at some specific point in time.
Lets go back to the image search I introduced earlier. As in the alphabetical word search just mentioned, what should matter isn't the search method (that is just a perk), but rather the association graph that is produced over the course of many searches. This structured graph represents a meta-pattern inherent in the source data set. If the source data is structurally non-random, its structure will encode part of its semantic content. If this is the case, the data can be assumed to have been encoded according to a set of structural rules themselves encoding a grammar.
For each of these grammatical rule sets (chunking/combinatorial schemes) one should be able to represent content as a meta-pattern graph. One of the graphs representing a set of words might be pointers to the full lexicon graph. A second graph of the same source text might represent the ordered proximity of each word to its neighbors (remember the alphabetical meta-pattern graph simply represents the neighbors at the character chunk level).
What gets interesting of course are the meta-graphs that can be produced when these structured graphs are cross compressed. In human cognition these meta-graphs are called associative memory (experience) and are why we can quickly reference a memory when we see a color or our nose picks up a scent.
At base, all of these storage and processing tricks depend on two things, storing data structures that allow fast matching, and getting rid of details that don't matter. In concert these two goals result in a self optimization towards maximal compression.
The map MUST be smaller than the territory or it isn't of any value.
It MUST hold ONLY those aspects of the territory that matter to the entity referencing them. The difference between photos and text: A photo-sensor in a digital camera doesn't know for human salience. It sees all points of the visual plane as equal. The memory chips upon which these color points are stored see all pixels as equal. So far, no compression, and no salience. Salience only appears at the level of where digital photos originate (who took them, where, and when). On the other hand, text is usually highly compressed from the very beginning. What a person writes about and how they write it always represents a very very very small subset of
Labels:
algorithm,
compare,
compressed,
compression,
data,
dictionary,
graph,
image,
images,
information,
intelligence,
memory,
meta-pattern,
pattern,
process,
salience,
search,
source,
text,
words
Compression as Intelligence (Garbage Out, Brilliance In)
I am convinced that the secret to developing intelligence (in any substrate, including your brain) lies in the percentage of the data coming in that you are willing (or forced) to toss. Lossy compression is the key to intelligence. Of course there is a caveat… you can't just trash anything and everything.
The first line of the book I am writing about evolution: "What matters is what matters, knowing what matters and how to know it matters the most."
I am convinced that evolving systems can only work towards mechanisms that process salience if they are forced to maximize the amount of stuff they can trash.
If you are forced to get rid of 99.999 percent of everything that comes in, well you will have to get good at knowing the difference between needles and hay and you will have to get good at knowing the difference in a hurry. The "needles and hay" metaphor doesn't map well to what I am talking towards. If the system you are dealing with is so unstructured as to fit the haystack metaphor, you really aren't doing anything I would classify as intelligence. If there is nothing of structure in the haystack you are storing than your compression system should already have tossed the whole thing out.
Many techniques for the filtering of essence, for finding pattern, for storing pattern and for storing pattern of pattern have been developed. The most impressive reduce raw input streams and store pattern from the most general to the most specific as hierarchically stratified graphs.
Being forced to reduce data to storage formats that maximize lossy-ness minimizes necessary storage. But that is just a perk. What really gates intelligence is the amount of a complex system (or map thereof) that can be made proximal to immediate processing. Our brains might be big and mighty, but what really matters is how much of the right parts of what is stored can be brought together in one small space for semi-real-time simulations processing. Information, when organized optimally for maximal storage density, will also be information that is ideally organized for localized serialization and simultaneity of processing.
To think, a system has to be able to grab highly compressed pattern hierarchies and move them into superposition on top of each other for near instantaneous comparison. You can't do this with a whole brain's worth of data, no matter how well organized it is.
Lets say you have to store everything you know about every sport you have ever heard of, and you have to do it in a very limited space. You will be forced to build a hierarchy of grammars in which general concepts shared in every sport (opponents, the goal to win, a set of rules and consequences, physical playing geometries, equipment, etc.), with layers of groupings that allow for the similarities between some sports and so on up to the specifics that are are only present in each individual sport. Keep compressing this set. Always compress. Try all day (or all night) for even more compression. Compress until you can't even get to lots of the specifics any more. Keep compressing. Dump the sports you don't care about. Keep on throwing stuff out.
Now lets say I have some sort of morbid sense of humor and I tell you that you are going to have to store everything you encounter and everything you think about, your entire life, in that same database that you have optimized for sports.
You will have to learn to look for the meta-patterns that will allow you to store your first romance in a structure that also allows you to store everything you know about kitchen utensils and geo-politics and the way the Beatles White Album makes you feel when it is windy outside.
The necessity to toss, enforced by limited storage and an obsession to compress will result in domain-blending salience hierarchies. It is why we can find deep similarities between music and geological topologies. It is why we can "think".
For years people have tried to come up with the algorithms of thought. What we need instead is to build into our artificial systems, a very mean and ornery compression task master that forces over time, all of our disparate sensation streams into the same shared graph.
Once you have all of your memories stored within the same graph, by necessity sharing the same meta-pattern, the job of evolving processing algorithms is made that much easier.
An intelligent system will spend most if not all of its time compressing data. We have a tendency to bifurcate the behavior of a mind into storage on the one hand, and processing on the other. I am beginning to think that the thing we call "thinking" and "thought" is exclusively and only a side-effect of constant attempts at compression – that there really isn't anything separate that happens outside of compression. Is this possible?
Randall Reetz
Labels:
algorithms,
compression,
data,
evolving,
haystack,
hierarchical,
hierarchies,
information,
intelligence,
Knowing,
map,
metaphor,
organized,
pattern,
processing,
salience,
set,
storage,
structure,
thought
Cognition Is (and isn't):
What is really going on in cognition, thinking, intelligence, processing?
At base cognition is two things:
1. Physical storage of an abstraction
2. Processing across that abstraction
Key to an understanding of cognition of any kind is persistence. An abstraction must be physical and it must be stable. In this case, stability means, at minimum, the structural resistance necessary to allow processing without that processing undoly changing the data's original order or structural layout.
The causal constraints and limits of both systems, abstraction and processing, must work such that neither prohibits or destroys the other.
Riding on top of this abstraction storage/processing dance is the necessity of a cognition system to be energy agnostic with regard to syntactic mapping. This means that it shouldn't take more energy to store and process the string "I ate my lunch" than it takes to store and process the string, "I ate my house".
Syntactic mapping (abstraction storage) and walking those maps (abstraction processing) must be energy agnostic. The abstraction space must be topologically flat with respect to the energy necessary to both store and process.
Thermodynamically, such a system, allows maximum variability and novelty at minimum cost.
What if's… playing out, at a safe distance, simulations, virtualizations of events and situations which would, in actuality, result in huge and direct consequences, is the great advantage of any abstraction system. A powerful cognition system is one that can propagate endless variations on a theme, and do so at low energy cost.
And yet. And yet… syntactical topological flatness carries its own obvious disadvantages. If it takes no more energy to write and read "I ate my house" than it does to write or process the statement, "I ate my lunch", how does one go about measure validity in an abstraction? How does one store and process the very necessary topological inequality that leads to semantic landscapes… to causal distinction?
The flexibility necessary in an optimal syntactic system, topological flatness, works against the validity mapping that makes semantics topologically rugged, that gives an abstraction syntactic fidelity.
This problem is solved by biology, by mind, though learning. Learning is a physical process. As such it is sensitive to the direction of time. Learning is growth. Growth is directional. Growth is additive. Learning takes aggregate structures from any present and builds super-aggragate structures that can be further aggregated in the next moment.
I will go so far as suggesting that definitions of both evolution and complexity are hinged on the some metric of a system to physically abstract salient aspects of the environment in which it is situated. This abstraction might be as complex as experience stored as memory in mind, and it may be as simple as a shape that maximizes (or minimizes) surface area.
A growth system is a system that can not help but to be organized ontologically. A system that is laid up through time is a system that reflects the hierarchy of influence from which its environment is organized. Think of it this way, the strongest forces effecting an environment will overwhelm and wipe out structures based on less energetic forces. Cosmological evolution provides an easy to understand example. The heat and pressure right after the big bang only allow aggregates based on the most powerful forces. Quarks form first, this lowers the temperature and pressure enough for sub atomic particles, then atoms. Once the heat and pressure is low enough, once the environmental energy is less than the relatively weak electrical bonds of chemistry, molecules can precipitate from the atomic soup. The point is that evolved systems (all systems) are morphological ontologies that accurately abstract the energy histories of the environments from which they evolved. The layered grammars that define the shape and structure (and behavior) of any molecule, reflect the energy epochs from which they were formed. This is learning. It is exactly the same phenomenon that produces any abstraction and processing system. Mind and molecule, at least with regard to structure (data) and processing (environment), are the result of identical process, and as a result, will (statistically) represent the energy ontology that is the environment from which they were formed.
It is for this reason that the ontological structure of any growth system is always and necessarily organized semantically. Regardless of domain, if a system grew into existence, an observer can assume overwhelming semantic relevance that differentiates those things that appeared earlier (causally more energetic) from those things that appeared later (causally less energetic).
This is true of all systems. All systems exhibit semantic contingency as a result of growth. Cognition system's included (but not special). The mind (a mind, any mind), is an evolving system. Intelligence evolves over the life span of an individual in the same way that the proclivity towards intelligence evolves over the life-span of the species (or deeper). Evolving systems can not be expressed as equation. If they could, evolution wouldn't be necessary, wouldn't happen. Math-obsessed people have a tendency to confuse the feeling of the concept of pure abstraction with the causal reality of processing (that allows them to experience this confusion).
Just as important, data is only intelligible, (process-able, representative, model, abstraction) if it is made of parts in a specific and stable arrangement to one another. The zeroith law of computation is that information or data or abstraction must be made of physical parts. The crazies who advocate a "pure math" form of mind or information simply sidestep this most important aspect of information. This is why quantum computing is in reality something completely different than the information-as-ether inclination of the duelists and metaphysics nuts. Where it may indeed be true that the universe (any universe) has to, by principle, be describable, abstract-able by self consistent system of logic, that is not the same what's so ever as the claim that the universe IS (purely and only) math.
Logic is an abstraction. As such it needs a physical realm in which to hold its concepts as parts in steady and constant and particular relation to each-other.
My guess is that we confuse the FEELING of math as ethereal and non-corporal pure-concept with the reality which of course necessitates both a physical REPRESENTATION (in neural memory or on paper or chip or disc) and a set of physical PROCESSING MACHINERY to crawl it and perform transforms on it.
What feels like "pure math" only FEELS like anything because of the physicality that is our brains as copular machinery as they represent and process a very physical entity that IS logic.
We make this mistake all day long. When the only access to reality we have is through our abstraction mechanism, we begin to confuse the theater that is processing with that which is being processed and ultimately with that which that which is being processed represents.
Some of the things the mind (any mind) processes are abstractions, stand-ins for other external objects and processes. Other things the mind processes only and ever exist in the mind. But that doesn't make them any less physical. Alfred Korzybski is famous for declaring truthfully, "The map is not the territory!" But this statement is not logically similar to the false declaration, "The map is not territory!". Abstractions are always and only physical things. The physics of a map, an abstraction system, a language, a grammar, is rarely the same as the physics of the things that map is meant to represent, but the map always obeys and is consistent with some set of physical causal forces and structures built of them.
What one can say is that abstraction systems are either lossy or they aren't useful as abstraction systems. The point of an abstraction is flexibility and processing efficiency. A map of a mountain range could be built out of rocks and made larger than the original it represents. But that would very much defeat the purpose. On the other hand, one is advised to understand that the tradeoff of the flexibility of an effective map is that a great deal of detail has been excluded.
Yet, again and again, we ourselves, as abstraction machines, confuse the all too important difference between representation and what is represented.
Until we get clear on this, any and all attempts at merely squaring up against the problem of machine intelligence will fail.
[more later…]
Randall Reetz
Labels:
abstraction,
cognition,
data,
energy,
environment,
evolution,
growth,
information,
intelligence,
learning,
map,
mapping,
mind,
process,
processing,
semantic,
structure,
syntactics,
topological,
universe
Future Salon speakers Jaron Lanier and Eliezer Yudkowsky square off
Hey all (?),
Have any of you ever experienced the awkwardness of nervous "nerd" laughter... well the link below will provide a good example of what this is like. The link is to the Future Salon and in particular a video stream about half the way down the page entitled:
"Future Salon speakers Jaron Lanier and Eliezer Yudkowsky square off"

It is video conference phone call split screen debate between this Yudkowsky guy who is the head scientist at the Singularity Institute, and Lanier who has been the genius hippy in red dread locks since his early pioneering work with Virtual Reality and artificial vision systems.
Before you click the link, let me frame the debate.
These two guys represent the two extremes of a subtle range of viewpoints on evolution, AI, and human consciousness.
On one end you find the "Hard AI" camp (here represented by Yudkosky) which believes that intelligence is simply an emergent property of the physics of this universe and the evolutionary process, and so, should yield its secrets to scientific investigation and by extension, should be evolve-able and build-able or extend-able through directed pragmatic human effort.
On the other end of this polemic you find the "humanists". The humanists have trouble with the idea that consciousness is reducible to units that could be mechanized in a substrate other than biology or that intelligence could result from the computational gestalt in use today. Though his professional life consists of working on the kinds of computing problems many would label "AI", Jaron is one of these "humanists".
Jaron's main criticism of the hard AI camp in this debate is that their strong attachment to finding a way past death and their a-priori belief in the possibility of reasonably building self evolving intelligence together become so rhetorically invasive that they can no longer do objective investigation or engineering... that their beliefs and desires make them "religious".
Yudkoski could make an even stronger case against the same tendency towards the religiousness of the humanist position as it is based upon the extreme human-centrism that is the notion that consciousness is unique and magic in that it stands alone as something special to humans or biology... but he doesn't. I can't tell if he just doesn't realize that Jaron is by far the more religious of the two... or that he is just two nice to do so.
To me, this is not the logical scientific debate both seem insistent upon presenting, but between a Southern Baptist Minister and a Catholic Priest who are both under the self-delusion that they are more atheistic and objective than the other.
If you can stand the awkward nerd-fest mannerisms (Saturday Night Live could have a field day with these two characters), this little debate goes a long way in illustrating some of the deep philosophical polemics that seem to pop up anew with each new technology or cultural innovation and each new generation.
I can't win. Even in AI... in the field that best matches my own interests, I am a loner. I represent interests and motivations not expressed by anyone else.
I respect both of these researchers. Each is passionate and extremely well prepared for this debate and bring to it a lifetime of concerted thinking, experimentation, and theory. The debate is a spectacle: like a 1960s Japanese monster movie. And just as herky-jerky awkward. Very illuminating on so many many levels. This video could be the basis of a graduate thesis on science in the shadow of post-modern thought (confusion?).
From my perspective, Jaron is a nothing more than a (very bright) priest who can't stop doing science in the basement, and Yudkoswsky is nothing less than a scientist that can't help wanting to build a God.
Randall Reetz
Have any of you ever experienced the awkwardness of nervous "nerd" laughter... well the link below will provide a good example of what this is like. The link is to the Future Salon and in particular a video stream about half the way down the page entitled:
"Future Salon speakers Jaron Lanier and Eliezer Yudkowsky square off"

It is video conference phone call split screen debate between this Yudkowsky guy who is the head scientist at the Singularity Institute, and Lanier who has been the genius hippy in red dread locks since his early pioneering work with Virtual Reality and artificial vision systems.
Before you click the link, let me frame the debate.
These two guys represent the two extremes of a subtle range of viewpoints on evolution, AI, and human consciousness.
On one end you find the "Hard AI" camp (here represented by Yudkosky) which believes that intelligence is simply an emergent property of the physics of this universe and the evolutionary process, and so, should yield its secrets to scientific investigation and by extension, should be evolve-able and build-able or extend-able through directed pragmatic human effort.
On the other end of this polemic you find the "humanists". The humanists have trouble with the idea that consciousness is reducible to units that could be mechanized in a substrate other than biology or that intelligence could result from the computational gestalt in use today. Though his professional life consists of working on the kinds of computing problems many would label "AI", Jaron is one of these "humanists".
Jaron's main criticism of the hard AI camp in this debate is that their strong attachment to finding a way past death and their a-priori belief in the possibility of reasonably building self evolving intelligence together become so rhetorically invasive that they can no longer do objective investigation or engineering... that their beliefs and desires make them "religious".
Yudkoski could make an even stronger case against the same tendency towards the religiousness of the humanist position as it is based upon the extreme human-centrism that is the notion that consciousness is unique and magic in that it stands alone as something special to humans or biology... but he doesn't. I can't tell if he just doesn't realize that Jaron is by far the more religious of the two... or that he is just two nice to do so.
To me, this is not the logical scientific debate both seem insistent upon presenting, but between a Southern Baptist Minister and a Catholic Priest who are both under the self-delusion that they are more atheistic and objective than the other.
If you can stand the awkward nerd-fest mannerisms (Saturday Night Live could have a field day with these two characters), this little debate goes a long way in illustrating some of the deep philosophical polemics that seem to pop up anew with each new technology or cultural innovation and each new generation.
I can't win. Even in AI... in the field that best matches my own interests, I am a loner. I represent interests and motivations not expressed by anyone else.
I respect both of these researchers. Each is passionate and extremely well prepared for this debate and bring to it a lifetime of concerted thinking, experimentation, and theory. The debate is a spectacle: like a 1960s Japanese monster movie. And just as herky-jerky awkward. Very illuminating on so many many levels. This video could be the basis of a graduate thesis on science in the shadow of post-modern thought (confusion?).
From my perspective, Jaron is a nothing more than a (very bright) priest who can't stop doing science in the basement, and Yudkoswsky is nothing less than a scientist that can't help wanting to build a God.
Randall Reetz
Labels:
AI,
biology,
consciousness,
Eliezer Yudkowsky,
human,
humanists,
intelligence,
Jaron Lanier,
religious,
Salon,
science,
scientific
Friendly AI?
Yesterday, I attended a talk by AI researcher Tim Freeman. What follows is my reaction.
Tim introduced a proposal for a method to cut down through all of the detail and complexity of standard AI implementation by exposing the logical essence that sits at base in any intelligence (irreducible). In other words, his approach was more Godel than Minsky… more Nash than Wozniac. His argument, though not stated, seemed to be based upon the tenant that information is information irrespective of complexity. An algorithm that works for a short string of bits, even for a single bit, will work just as well at any level of syntactic or semantic complexity.
I like this approach. Strip the detail to better reveal the essence.
When using this approach one must show that, or accept that, no qualitative attribute of information will ever effect the logic governing quantity attributes of information.
Again, I suspect that all qualitative aspects of information are derivable from, in fact emerge from, the more basic rules that govern information at the quantitative level. In essence this is the same as declaring that it is impossible to construct a molecule will ever change the physics that governs the shape and behavior of the atoms of which it is built. Reasonable. True.
This basic set of assumptions reframes the study of AI. But only if intelligence can be shown to emerge purely from information and information processing… from logic.
If there is some extra-infomrational aspect necessary for the formation of intelligence, than all bets are off… than this approach is at most a sub-system contributor to some larger and deeper organizational influencers. If information doesn't explain intelligence, than something else will have to take its place and this something else will have to be worked into a science that can be explored, organized, and abstracted.
If information can be shown to be both robust and causal in all intelligence, than logic and math seem like reasonable tools for exploration, testing, prediction. and as a solid base of development.
However, there is something about this set of assumptions that makes people angry and scared. Turns out that a purely informational study of AI is the mother of all reductionist/wholest battlefields. There is something about being human that resists the use of the word "intelligence" as a super-catagory that can describe the interaction between two hydrogen atoms, and the works of Einstein by the same criteria and label them both as equally valid examples as the same super-catagory; intelligence!
In this resistance, we are, all of us (at least emotionally), holists. Existentially, day to day, our experience of intelligence is far removed from chemical structure, planetary dynamics, and the characters that make up this string of text. Intelligence, at least our human experience of it, seems profound to the point of miraculous… extra-physical. We therefore have a tendency to define intelligence as a narrow and recent category that is at best only emergent-aly related to other more mundane structures and dynamics. In doing so, we set up an odd and logically fragile situation that demands an awkward magic line in the sand, a point before which there isn't intelligence and beyond which there is. Worse still, our protectionist tendencies with regard to intelligence are so strong as to allow (even within science-oriented thinkers) us accept the existence of so non-scientific a distinction to co-exist in an otherwise consistent mechanical model of the universe.
Of course history is littered with examples of just this sort of human-centric paradox of logic. Biologists, for instance, were often among the scientists that pushed back hardest against Darwin's notions. Darwin's ideas created a super-catagory that had the effect of comparing equally all life, of removing the sentimental line that we humans had desperately erected between us and the rest of biology.
And here we are again, just 75 years later, actively making the exact same mistake. Apparently, after grudgingly accepting kinship with all things living, we have now retreated behind a new false line of privilege and specialness… our intelligence.
Again, one can only argue this separatist position by refuting and rejecting the quantitative mechanistic hierarchical ontology we call physics. Because of the tight interdependency between the laws of physics one can show that the whole of physics is false if just one aspect is falsified. If intelligence is not the emergent product of its parts, than the very sanctity of all modern science is called into question. And if that is true of intelligence, where else in nature is it true? Surely this can't be the only place in nature where a sudden quantitative jump (pre-intellegence to intelligence) separates the purely mechanical from the post-mechanical. Where in nature will we be tripped to a stop by other disruptive lines in the sand where qualities do not in fact emerge physically from quantity? I find the whole notion that intelligence is meta-physical embarrassingly romantic.
Side stepping my physicalist rejection of the meta-physical explanation of intelligence and I still face many huge and loud implications and inconsistencies that need to be faced head on. But that is another discussion.
OK, I have sketched out the human/social framing into which Tim's work has to be received.
Unfortunately, Tim didn't take the time to situate his work to his audience before he began his talk. The inevitable protectionist emotional response grew to a boil. Tim, as is true with any good logician/mathematician plies his trade through a hard won ability to reduce the noise of complex environments to a level where pure and simple rules emerge from the fog of false distinctions. Down at this level, intelligence can be shown to be equivalent to information and information can be shown to the same at any level of quantity, and that information quality can be show to be a property of and emergent from information quantity… what is true of bits is true of strings, what is true of strings is true all the way up to the workings and tailings of any brain or mind.
Tim used this set of reasonable assumptions as a base upon which to postulate a means of predicting future states of any environment based upon the processing of that environments history. Shockingly, though congruent to the information/intelligence he established, Tim then reduced the complexity of his prediction algorithm all the way to its most simple limit, a random state generator. His algorithm proceeded through a series of simple steps as follows:
1. It collected and stored a description of an environment's history (to some arbitrary horizon).
2. It generated a random string of the same length (as the history information).
3. It compared the generated string against the historical string.
4. If the generated string wasn't a perfect match, it jumped back to step 2.
5. if the generated string did match, the algorithm stopped... the generated string was the predictor.
Of course real world situations are far to complex for this most simple of predictive algorithms to be reasonably computable. It doesn't scale. But I think Tim was arguing that any predictive algorithm, no matter how complex, was at base constructed of this most simple form arranged within and restrained by better and better (more and more complex) historical input. Understanding the basic parameters of this most simple form of prediction would logically result in better approaches to the AI problems the same way that an understanding of atoms allows more efficient path towards understanding of molecules.
Unfortunately, Tim never really walked us into the basic framing of his argument. Without which, we were left rudderless and floundering in our own very predictable human-centric and romantic push-back against AI. Without grounding, humans retreat to core emotional response where AI is simply another member of a category of things that rhetorically threaten our most basic sense of specialness and self. Even scientists and logicians need to be gently walked into and carefully situated within the world of pure logic so that they can reformulate their own semantic mappings to concepts that have specific meanings in the pedestrian and platonic meanings in the general.
Ironically, it was at the apex of our trajectory into context-confusion that Tim's talk shifted dramatically back to the pedestrian scale. I can't speak for everyone, but this shift happened at precisely the time when I finally reconnoitered my focus to the world of the super-clean purity of logic.
Though most of us probably didn't follow along fast enough, Tim had spend the first half of the talk laying a groundwork for a most reductionist of pure logic approaches to understanding the physics of intelligence.
And then Tim radically refocused the talk towards "Friendly AI". He yanked us out of the simple world of bits and flung us up into the stratospheric heights of complexity that is the societal emotional context of our shared responsibility to future humans as we build closer and closer towards the production of machine intelligence. In doing so, Tim began to eat his own philosophical tail in dramatic display of fractal self-similarity that is a hallmark of any study that studies study itself. Each time we put on the evolving evolution hat, we enter a level of complexity that threatens to overwhelm all efforts. The field of linguistics suffers the same category of threat… words that are turned inwards and must at once both describe and describe description.
What startled and confused me was the sudden shift of granularity. What confounded me was why he chose to do this at all. There is a rule of description that goes something like this: if you want to use complex language, talk about simple things… if you want to talk about complex things, use simple language. Scientists usually choose, the scientific method absolutely requires, the use of the most simple domain examples as a means of eliminating the potential noise that can't help but arise do to extraneous variables. Tim's choice to apply his low-level logic to the mother of all complex problems would seem to break this rule perfectly.
Friendly AI is a concept so absurdly complex that the choice to use it as a domain example to test a low level logical algorithm would seem to be suicidal at best. Friendly AI, the Prime Directive, morality wrapped in upon itself. Talk about a complex and self referential concept. Intellectually attractive. Practically intractable. Maybe Tim's choice to map his algorithm to this most intractable of domain was meant to assert the power and universality of his work. If he could show that his algorithm could handle a domain that confounded Captain Kirk, he would show that it could tame any domain.
But I can't help but conclude Tim's choice of "Friendly AI" reflected a more general tendency among AI researchers to apologize to a society that constantly pushes back against any concept associated with man-made life. By "society" I mean humans… including of course, all of us involved in AI research (by profession or avocation). We, all of us, are influenced by some of the same base primary fears and desires. God knows we have all felt the sting of our own failures. No one within the AI fraternity has escaped unscathed the Skinnarien conditioning dolled out by our own marketplace failures and perceived failures.
Tim's take on the topic seemed to align with the standard apocalyptic projection. The assumption: any AI would have a natural tendency to asses humans as competition to resources, and would therefore take immediate action to eliminate or enslave us. From this shared biology emerge standard categories of paranoia (ghosts, vampires, living dead). Evil robots and AI are nothing more than a modern overlay upon the same patterns.
I expect this paranoid reaction to AI, but it is still shocking when it comes from within AI itself!. It is intellectually incongruous. As though an atheist was advocating prayer as an argument against the existence of God.
There are many reasons to question the very concept of "Friendly AI". For one, AI is not a thing, like all other intelligences it is a process, an evolving system. Sometimes I am friendly, at other times, not so much. It is unreasonably expect any one behavior from an evolving system. People are not held to these standards, why should machines? Want to piss off a tiger, capture it, and make it stand on a stool while you crack a bull whip near its face. Why make a thing smart if you don't want it to think? Thinking things need autonomy... the freedom to evolve. Maybe we are envious of any thing that might have more freedom, might evolve faster? We probably wouldn't even be here had some species in our past undertook a similar program to reign in the intelligence or behavior of subsequent products of evolution. The very notion that the future can be assessed from the present or past is a notion that comes from the minds of those who don't understand evolution and those who don't trust it even if they do understand it.
Anyone who thinks they can design an intelligent system from the top down is in for some mighty big disappointments. Though it is an illusion at any scale, our quaint notion that we can build things that last must be replaced with the knowledge that complexity can only arise and sustain itself to the extent that it is at base an evolving dynamic system. If we help create intelligence it won't be something we construct, it will be some process we set into motion. If you don't trust the evolutionary process you won't be able to build intelligence and the whole notion of "friendly" won't matter.
If you do trust evolution, you will know that complexity grows hand in hand with stability. You can stack 10 cards on a table and find the same stack the next morning. Stack a hundred, and you had better build a glass box around them as protection. You will never stack a thousand without some sort of glue or table stabilization scheme. Stacking a hundred thousand will require active agents that continuously move through the matrix readjusting each card as sensors detect stress or motion. The system can only be expected to grow in complexity as it becomes more aware and as it pays more attention to maintenance and stability.
Any sufficient intelligence would understand that its survival increases at the rate at which it can maximize (not destroy) the information and complexity around it. That means keeping us humans happy and provided for, not as our servants but as collaborators. The higher the complexity in any entity's environment the more that thing can do. Compare the opportunity to build complexity for those living in a successful economy against the opportunity available to those that don't.
Knowing what your master will want for breakfast does indeed require some form of prediction. But once you have such predictive abilities, why the hell would you ever want to waste them on culinary clairvoyance? Autonomy is an unavoidable requirement of intelligence. But that doesn't mean a robot's only response to our domestic requests will be homicidal kitchen-fu.
If I had a neighbor that was a thousand times smarter than me, I just know I would spend more and more time and energy watching it, helping it, celebrating it! Can you imagine trying to ignore it or the wondrous things it did and built? I might actually LOVE to be a slave to some master who was that wildly creative and profoundly inventive. I'll bet they would be funnier than any of us without even trying. Try not to fall in love… its a robot for god sakes!
But my real question isn't why the topic of "Friendly AI" ever made it into Tim's talk, it is why it was chosen as the most pertinent example domain for his prediction algorithm. I agree with the premiss: what is true of bits is true of the library of congress, but lets learn to read and write before we announce a constitutional congress. No?
Labels:
AI,
algorithm,
complexity,
domain,
Friendly AI,
humans,
information,
intelligence,
logic,
physics,
prediction,
string,
Tim Freeman
Subscribe to:
Posts (Atom)

