Showing posts with label coursework. Show all posts
Showing posts with label coursework. Show all posts

Thursday, May 3, 2012

Games, Language, and Meaning

Note: This is my second (and final) paper for my cognitive science survey course. It is a response to Robin Clark's Meaningful Games. This is ~4300 words including lots of armchair prognosticating about language and rhetoric as well as plenty of complaining about the ideas Clark presents in his book. You have been warned.

Within five minutes of Professor Eisenberg mentioning Meaningful Games in class, I had placed it in my Amazon shopping cart. I am not certain I had ever been more excited about a book. I am positive that a book for class had never excited me as much as the idea of this one did. You see, Robin Clark’s book proposes to “explore language with game theory”, an idea that immediately intrigued me as a reasonable, if not immediately apparent, approach to a topic that we tend to treat with a highly analytic eye despite its undeniably human nature.

I will say right now that I see no good reason to look at language as something that’s biologically evolved or that can be universally described. We spent time in class looking at animal communication that we described as being somehow sub-linguistic. We noted that many species, especially more evolutionarily advanced ones such as primates, seemed to have developed distinct cries for denoting various circumstances in the environment -- leopards, snakes, etc. The tendency to vocalize and share information with the surrounding world, then, would seem to be evolved. I do not believe that language (the fusion of semantics and syntax) itself is, though.

Language seems, to me, to be more like a tool. The basic components of this tool -- those surprisingly varied vocal notifications that animals of many advanced species use -- are available to less evolved creatures. They use them in their simple way, just like otters cracking open sea urchins with rocks and chimpanzees fishing for termites with sticks. But humans are more advanced. We perceive more of the world and recognize relationships in the world more acutely. Just like we combine simpler physical resources into more elaborate contraptions, why wouldn’t we devise ways of combining vocal emissions to more effectively and efficiently relate truths about our surroundings -- in both the immediate and abstract senses. Would anybody say that (specific) tool use is evolved?

Would anybody claim that Japanese macaques have somehow become genetically predisposed towards washing their sweet potatoes? That seems downright silly. Why, then, would I have a genetic determination to recognize noun phrases and formulate clauses? Some people may point to the similar structures of languages across the globe as a suggestion that we are built to process languages in a specific way. I would say that even these similarities can be explained by the tool metaphor.

First of all, tool use is taught generationally. Successive generations, being taught how to make and use the tools of their forebears, do not have to waste their lives re-inventing the basics. Instead, they can move on to expand on existing technologies or invent new technologies -- including linguistic “technologies” -- previously unimagined. Over time, some tools survive and others are replaced. Some tools are simply better than others, and people with superior technologies of the moment tend to spread farther and subjugate those with weaker tools. In this way, the technological ecosystem evolves concurrently with our culture.

Linguistically, we expand and evolve our tools by inventing new ways of describing new experiences and more effective ways of describing old ones. Today we see this primarily through the invention of words, but that’s the microevolution of a language. It’s quick and dirty and gets the job done at the moment. Is it unreasonable to think of the addition, removal, and adaptation of grammatical structures as the macroevolution of a language? A more flexible syntactic structure is going to allow for more efficient communication, and thus ease education and cultural advancement in other areas. Just as multiple civilizations concurrently developed nautical and agrarian technologies that eventually converged to similar solutions, it’s reasonable to believe that grammatical structures across independently developed languages would also converge to similar standards.

Why would these structures be similar? Because they need to be processed in the brain. We live in a physical world, and the features of this world, including our biology, can be modeled mathematically. As such, there are logical structures, perhaps nested hierarchies like those often used to describe modern languages, that may simply coordinate better with the design of our brains. Just like we wouldn’t build a scythe that takes three arms to wield, we wouldn’t keep around a language for long if it did not make concessions to the way we focus our attention and process information. Over time, it’s very possible that we could select for traits that allow individuals to use an especially important tool more efficiently. Certainly, I operate in a mental space that is almost entirely linguistic, but I still see no reason to think that language itself is built into us any more than a tendency to build and drive automobiles.

I apologize for this protracted rant, but it’s necessary to explain my excitement over the prospect of Meaningful Games. If language is, as I see it, a socially constructed tool and not some sort of Platonic Form that has always existed, then it makes little sense to analyze language as something independent of communication. Communication fundamentally occurs between rational agents, and rational agents tend to be aware of one another’s rationality. When I hear the words “exploring language with game theory”, I presuppose an account based on this principle. I think of explanation that would account for the semantic richness of and the syntactic flexibility of languages from a perspective of the ways we use language; a perspective that respects the fundamentally social nature of the topic.

Clark’s preface to Meaningful Games suited to wet my appetite quickly. He touches on some of these very same issues I mention, especially the communicative purpose of language. He recognizes that language is defined by its use and that its use is defined by social circumstance. He talks about Tarski and Grice and logical models and meaning; he talks about truth making and algebras and grammars and what these things mean in a context of multiple rational agents. He talks about social context and developing techniques to describe language by its use -- to evaluate it, essentially, as a tool. He said all of the right words to get me excited about what he had to say.

Clark then begins his book proper by discussing the concept of “mentalese”, the sort of universal, Platonic “language” of thought. Mentalese, as Clark describes it, is basically a set of constructs that describe the interactions of objects and agents in the world as we perceive it. These constructs are expressed in a sort of predicate form, relating state transitions on individual objects and causal relationships between these transitions. For example, Clark gets hung up on the notion of what it might mean for A to “kill” B. He denotes a predicate kill(A,B), which suggests that A caused B to die.

We can imagine a complicated series of such predicates and the potential to conjoin and disjoin them and forming propositions of them in first order logic to express even more complicated relationships. These predicates can then be compared to the world we know and we retain those predicate pairings we know to be true. We can even conceive of what a world would be like if other predicates held true, as well.

The problem that Clark has with Mentalese is fairly simple: as Clark describes it, Mentalese effectively embraces the computational metaphor. Mentalese is an analytic thought form, only serving to hold a representation of the universe that is dependant on a detached, sensory experience that may or may not accurately resemble the world as it truly exists. Verbal language, under the assumption of mentalese, is simply a way of translating our own mentalese into an intermediate language that another person can then place back into their own version of mentalese. But this has two problems.

First, of all, we obviously do not operate simply on first order logic. We regularly use constructs like “most” or “many” or “often” that are vague and difficult, if not impossible, to construct into first order logic forms. But first order logic is the most complicated logical system for which we have correct and complete proof systems for. We cannot just be computers, then, because we cannot simply compute much of what we store and communicate. More troublesome for Clark is the simple fact that a world where we only work in mentalese is a solipsistic one, where individual understandings of reality are isolated, answering only to some higher standard. In the real world, though, it appears as though meaning is derived socially.

Clark lays these objections out in such a convoluted form that I am not sure either of us understands where his problem with Mentalese actually lies. He walks through the uncanny valleys of the Turing Test and the Chinese Room en route to an objection that is social rather than computational. And he never quite seems to resolve the two issues. Clark spends a chapter examining the social nature of meaning before taking time to describe game theory and the way we evaluate the design of games and strategies. By the time he actually starts to talk about linguistic games, suddenly people are computers once more, though he may not realize it.

Part of the problem is that the first linguistic “game” that Clark constructs is not really a game at all. It involves the verification or falsification of statements, the very thing Clark suggests Mentalese is designed to do. (In)Appropriately enough, Clark spends an extended amount of time discussing how we might translate statements into propositions in first order logic and how players might take turns stripping down the logical operators to check the truth of the statement against some publicly shared model of the world.

After nearly 60 pages, Clark acknowledges that this is a trivial game and not particularly useful. What he does not acknowledge is that his “game” fails to be a game. There is no actual strategy to be played, no matter what Clark might suggest. A proposition in first order logic is either true or false with respect to a model. The “winner” is determined not by player behavior but by the model over which they are playing. In the middle of all of this, Clark establishes more troublesome trends that will persist throughout the rest of his games. For one thing, his players are not communicating with each other. They are projecting information into and extracting information from a void. More important, none of Clark’s players behave in any less computational a manner than those in his verification game.

These problems become most clearly apparent when Clark starts talking about common knowledge. This chapter begins with an example of generals trying to coordinate an attack on a superior force. Not wanting to risk attacking alone, they send messengers back and forth between camps ad infinitum, believing that they cannot be certain that a confirmation was received until they receive a confirmation in turn.

Clark extrapolates this problem to individuals trying to discuss going to see a dance troupe when one performance has been cancelled and replaced with another. The problem, in this case, revolves around the definite description “the dance troupe” as opposed to specifying which troupe by a unique name. In our fantasy land, both players know about the cancellation, but neither knows what the other knows. They want to avoid confusion, but they also want to save effort of using the full name for the troupe that is actually performing. Clark suggests that this requires that they model each other’s mental states. That is, A needs to know about what B thinks A knows. Which means A needs to know about what B thinks A thinks B knows, and so on and so forth.

For asynchronous (or mono-directional) communication, this is a reasonable (if absurd) problem. We need to be aware of what our audience may not know and make concessions to ensure clarity. Direct communication between individuals, however, is inherently synchronized. When conversing, we may attempt to model one another’s knowledge, but we know that we can count on our partner’s awareness of their own knowledge to ease the burden. Individually, we can recognize when uncertainty enters the picture and make explicit reference to it. We can request clarification when we encounter confusion. Such active error correction may require more effort than simply knowing what referents are in play during a discussion, but it certainly requires less effort than infinite recursion.

This is what I mean when I say that Clark, despite his earlier complaints, continues to treat people like computers. We are communal creatures who operate on heuristics. We share information, including information about our own knowledge. Moreover, we generally do not build exact models or find exact solutions to problems as our first course of action. Rather, precision is a last resort once our estimations prove insufficient. Why would Clark even waste space talking about an effort to build a perfect representation of the cognitive world of others? This is not something we even attempt in indirect communication, as heuristics of charity encourage greater clarity than may be necessary. No model is perfect. The only perfect map of a territory is the territory itself. Anything else, by necessity, involves abstractions meant to communicate only the most significant details at the appropriate scale.

It is at this point that I started to get frustrated with Clark’s writing and began wondering if and when we would get to the meat of things. Was there ever a point where we would look at “games” that were concerned with something more elaborate than ambiguity? So I read. And I read. And I read some more. Eventually, on page 197, when Clark was still discussing types of copresence and their usefulness in resolving ambiguity in pronouns and definite phrases, he provided these two sentences:

(19) Because it was broken, I returned the plate I had just bought to the store
(20)Because the plate I had just bought was broken, I returned it to the store

Clark’s response to the variation in this sentence is as follows:
“The pairing of (19) and (20) immediately suggests another level of strategic decision making, one more allied to the general problem of stylistics and rhetoric. I leave this as an open problem for the reader.”

Suddenly everything that led to this point made sense. Clark is not concerned with communication strategies in the same way I am. When I think about use of language between agents, my concern, first and foremost, is with the rhetoric that Clark ignores. I am intrigued by strategies we use in our diction and our phrasing to affect the way our words are perceived by others. I am curious about implicature and subtler ways we make our ideas more or less palatable to listeners, and the ways in which we listen for these cues in an effort to reduce their impact.

Clark, on the other hand, is solely concerned with notions of clarity and economy. He looks at meaning in explicit ways, ignoring implicature in favor of understanding how sentence structure and “public information” (the very existence of which, Clark questions) can alleviate ambiguity. He invents quantifications of linguistic economy in an effort to explain our use of ambiguous phrases. He then justifies these efforts by making games to describe how we spare ourselves effort in our speaking and the risks that we take in doing so. Even in this regard, on a playing field of Clark’s own creation, I am not entirely certain how his “games” hold up.

Let’s look at an example involving pronoun use. We start out with two players, the speaker and the listener. We then have a statement. Clark uses the example “An undercover cop was observing a suspect. He stayed in the shadows.” In order to eliminate lexical cues, we could generically use “X ____ed Y. He ____ed.” In any case, Clark looks at this game as one of information states. The speaker begins in one information state, in which “he” refers to either X or Y, depending on the situation. His goal is for the listener to reach the same state. The speaker then has four options. He can use a proper name in place of “he”, or he can use “he” to refer to whichever of X or Y he intends for the listener to identify with the pronoun.

If the speaker chooses a proper name, then the listener has an easy choice: she can recognize the subject of the second sentence as the one directly referenced. Communication succeeds, both parties benefit. If the speaker chooses to use a pronoun, the listener gets to choose which of X and Y she believes the speaker to be referring to. A failure to choose the proper antecedent leads to confusion and both parties lose. Choosing the correct antecedent results in a benefit to both parties.

Obviously, to make this a game worth playing, the benefit for successful communication using a pronoun needs to surpass the benefit gained from using a proper name. As mentioned earlier, Clark is concerned with a concept of “economy” in language use. Since it takes more effort to generate a proper name (or a more descriptive phrase) than it does a simple pronoun, it is beneficial to the speaker to use the pronoun. Supposedly, it is also easier for the listener to process the shorter phrase. With this in mind, he devises a game that effectively looks like this, with the speaker’s options listed along the top and the listener’s options listed along the side:


Proper reference to X
“He” referring to X
“He” referring to Y
Proper reference to Y
X
(1, 1)
(3.5, 3.5)
(-2, -2)
(-2, -2)
Y
(-2, -2)
(-2, -2)
(2, 2)
(1.5, 1.5)

Where do these values come from? Imaginationland, of course! That’s not quite fair. Clark does have some rules that he follows to create the rankings. Successful communication, obviously, carries the highest benefit and unsuccessful communication the highest cost. Encoding an element as a pronoun provides a larger benefit as opposed to using a full name. More prominent elements (such as the subject of an immediately previous sentence) are “cheaper” to encode, increasing the benefit of that encoding. Not immediately relevant is the fact that pronouns are cheaper than descriptions.

Continuing to ignore the fact that the actual values are wholly arbitrary, this is why the highest reward comes from using “he” when the speaker wants to refer to X. Playing the game backwards, we see that this is the choice the speaker should make if, indeed, he wishes to refer to X in the second sentence. If the listener hears “he”, she will be inclined to assume that there is a 50% chance that the pronoun could be in reference to either X or Y. X’s prior prominence makes it the safer option. As such, if the speaker wants to refer to Y in the second sentence, he ought to make the reference explicit.

Does this final decision make sense? Mostly, again assuming that the context fails to provide further lexical cues. What doesn’t make sense is the reasoning that led to it. Specifically, why is the pronoun beneficial to the listener? She has to take the mental effort not only to process the auditory emission, but to attach it to some other concept. Depending on the situation, this could be a lot more work than processing an exact name. Pronoun use is not symmetrically beneficial. It is selfish on the part of the speaker.

Of course, this is not the only game that Clark plays. He looks at situations where individual words create ambiguity by carrying multiple definitions. He also looks very lightly at implicature through the lense of sarcasm and hinting/politeness. In these cases, he suggests that we can design games that use social context, familiarity, and/or focal points to create probabilistic strategies. Prior experience as well as physical and lexical context dictate our initial response to a word’s definition when there are multiple options. Familiarity determines our likelihood at properly interpreting sarcasm or recognizing a hint as such.

Of course, with every situation being unique, we can’t devise universal games for these sorts of problems any more than we can for pronoun use. But Clark thinks that the approach is still useful in the way that it grounds our use of language in a particular context. I do not. I do not believe that games are at all a necessary construct in recognizing how we generate meaning. I am not even sure this central question of meaning is a linguistic one. The bigger issue, though, is this: does anybody actually think this way when speaking?

Here we find the real problem with Clark’s attempt to leverage game theory on topics of meaning and ambiguity. Game theory requires that its participants are rational decision makers who are, at some level, aware of the game that they are playing. In an ordinary, conversational setting, such as those that Clark describes, I do not think that human beings can be described as rational actors. We tend towards preoccupation. If we are playing a game, it is one of minimal energy expenditure, and over thinking every lexical decision we make is certainly not a winning strategy in that situation.

In order for game theory to be applicable to a given situation, there needs to be a reason for maximization. In order for human beings to behave as rational actors, there needs to be meaningful cost or reward. There needs to be risk. Under these sorts of circumstances, we are more likely to take the time to explore our options and try to improve our expected result. Typically, conversation is not such a situation. We are free to identify and correct misunderstandings. Any serious consequences are not immediate and irrevocable, and any direct consequences of poor communication are not overly costly.

If we are looking for a situation where people play games with language, designing (and responding to) rhetoric is exactly such a situation. When we engage in rhetoric, as a writer or an orator, we work with an agenda. We manipulate language as a tool to promote that agenda, and our ability to use language and anticipate the response of our audience is what determines immediately our success or failure. As an audience, we try to identify linguistic ploys and focus on facts. The game is plain, but, as an example, I suggest my experiences serving jury duty.

In the summer of 2010, shortly before moving to Colorado, I was under going voir dire in the Hennepin County court system. Since I was preparing to leave the state and had a lot of preparation to do, I did not want to actually sit on a jury. Effectively, the attorneys and I were locked in a game. They asked questions. trying not to violate procedure or give away the case they were building but still seeking jurors who would be sympathetic to that case. I phrased my answers in such a way that, without lying or being contemptuous, I could suggest that I would not be easy to sway or take kindly to their efforts at persuasion.

Both of us had something to gain and/or lose. I might have to risk days of my life hearing a case I really didn’t have time to hear, and possibly a boring one at that. The attorneys were trying to find people who would be willing to consider the story they wanted to weave. We played the game with language, leveraging various grammatical forms and semantic particularities to implant opinions and ideas without violating the constraints of the situation. By acknowledging and responding to their rhetoric, I pushed the situation towards one that was favorable to myself -- and which I believe was favorable for the attorneys.

So, rhetoric is obviously a game, and likely a better game than the ones that Clark creates in Meaningful Games. As much as this helps explain my frustrations, the fact that I would have preferred this as a topic has no bearing on the quality of the book itself. Even discounting these desires, though, Clark’s approach and his games did not do much to think about the way I think about language or say anything novel about meaning. Maybe I would feel differently if I were a linguist, but Meaningful Games mostly just left me frustrated with Clark’s insistence on treating heuristic bits of human behavior with such an analytic eye.

By my estimation, Clark seriously inflates the value placed on economy of speech. The best reason to do so is simply that it lends credence to his arguments, but it also means he overcomplicates the entire issue. There is also the problem of incessant tangents -- about mentalese, garden path sentences, focal points, logical proofs, and anything else Clark thinks of as mildly relevant. While fascinating in subject, the added value is of little comfort when the stated purpose of the book comes to so little.

Had I entered with lower expectations, I may have enjoyed Meaningful Games. Clarks’ writing is, for the most part, accessible, and his examples tend to be light and entertaining. It takes a joy in itself and manages to cover large swathes of ground, even if much of it is tangential. I would actually consider recommending Meaningful Games to people with no prior knowledge of any of the topics contained within it, but certainly not to a person looking for a serious read on topics of game theory, language, or meaning.

Work Cited:
Clark, Robin. Meaningful Games: Exploring Language with Game Theory. Camebridge, Mass: The MIT Press, 2012

Thursday, March 8, 2012

Computers, the Cortex, Prediction, and Intelligence

**Warning: This is a response to the book On Intelligence by Jeff Hawkins. It was originally written for a Cognitive Science course. Ahead lies 4000 words worth of historical computer science, bumbling neurobiology, and a bit of armchair philosophy. Read at your own peril**

Jeff Hawkins, author of On Intelligence has a bone to pick with both Artificial Intelligence researchers and neuroscientists: neither group, he claims, seems to be concerned with determining the nature of intelligence. For their part, computer scientists have long been resistant to the notion that the structure of the brain is important in promoting intelligent behavior. Neuroscientists, meanwhile, have not put forth as much energy as Hawkins would like into putting forth any sort of unified theory or framework of cognitive function.

The sins Hawkins accuses computer scientists of committing can seemingly be attributed to an over-reverence for Alan Turing. If it is true that a Turing Machine can compute anything that can be computed, and if any we believe that one of the primary functions of the brain is to compute, then it seems perfectly plausible to suggest that the brain is nothing more than a biological implementation of a Turing Machine. Since all Universal Turing Machines are effectively equivalent, then it further seems reasonable to insist that digital computers ought to be a perfectly suited to playing host to intelligence.

Hawkins also argues that Artificial Intelligence researchers were led astray from the very beginning by the Turing Test. Turing lived at a time when psychology was still dominated by behaviorism: the notion that intelligence could only be determined by action. The Turing Test endorses this thinking, and while it is not a benchmark that researchers seriously pursue, virtually all testing of machine intelligence that has followed in its wake is also focused on this input-output driven benchmarking.

The largely fruitless results of historical AI research suggest that maybe this is not the right way to do things. We have made programs and algorithms that can do all manner of seemingly complicated activities from playing chess to approximating optimized configurations for complex systems. The “simple” things that we would really like computers to do, though (vision, language acquisition, motor control) have made only minimal progress. If the brain is following formal, procedural algorithms in the fashion of a turing machine, we obviously have not found them. More likely, though, is that the brain does something different.

Hawkins suggests that the brain, while obviously capable of traditional computation, must also do something more elaborate, and we ought to try replicating that process if we want to create computers and/or programs that exhibit intelligence. Computer Scientists came up with this same idea decades ago and it led to the idea of Neural Networks. This was a good start, but it was not taken far enough. As soon as toy three-layer networks produced some interesting results, researchers returned to navel-gazing instead of taking further steps to mimic brain behavior. But who can blame them when we still understand so little about the brain?

This leads us to Hawkins’ frustrations with the neuroscience community. Chiefly, he thinks that we simply have not put in enough effort to determine how the brain, and specifically the cortex, works. We have gathered reams of experimental data about brain activity. We have rough maps of where all sorts of phenomena are processed in the brain -- from linguistic syntax to motor control to immediate optical stimulus. What we do not know is really anything concrete about how these signals interact to produce what we think of as intelligence.

This is not really surprising in itself. The brain is an incredibly complex and sensitive organ, and it is very hard to perform any sort of accurate experimentation on it. All of our current methodologies necessarily sacrifice either spatial or temporal resolution, and we really need both in order to say anything meaningful. Still, Hawkins would encourage boldness. We cannot spend all of our time simply collecting data without anything to use that data for, and our current theories are too low-level and low-risk to be of interest. We need a theory that can explain the phenomenon of intelligence as a whole, both in order to understand ourselves and in order to imbue this property on future generations of machines.

To this end, Hawkins suggests an overarching theory to describe the nature and purpose of brain function as it relates to intelligence. Such a theory, while likely flawed, would give direction to our research and change the way we think about our work, both in neuroscience and A.I.. Ultimately, On Intelligence is Hawkins’ attempt to lay out such a theory: the Memory Prediction Framework.

In short, the Memory Prediction Framework suggests this: The function of the brain, as leads to intelligence, is not simple computation. Computation would suggest nothing more than stimulus-response pairings. Instead, Hawkins claims that the human brain strives to use previously identified patterns in order to predict and change the future.

For Hawkins, the seat of what we think of as intelligence -- agency, intentionality, adaptability, even creativity and consciousness -- is the cortex. The cortex is not exclusive to humans, but the most notable difference in brain structure between us and other mammals is the sheer size of ours, thanks to the evolutionarily recent expansion known as the neocortex. Since humans are so far beyond other animals in our ability to understand and control our environments and pass on our knowledge to subsequent generations, this enlarged cortex ought to play the primary role.

Hawkins proposes that the chief function of the cortex is to contain a model of the world as we have experienced. At its most abstract, ignoring all of the biology involved, this is accomplished by forming what Hawkins calls invariant memories. We recognize patterns that occur together -- the shape of a hand, the sensation of heat on our skin, the sound of a musical interval -- and group the sensations together into a single mental concept, recognizable even when specifics -- the starting pitch, the orientation of the hand, the location of the burn -- change.

From these patterns, we construct composites of patterns. From letters we pick up words and phrases and stories. Eyes and nose and lips become a face. Intervals expand to phrases combine to make songs. Simple patterns, universally, become building blocks for more complicated concepts, both in the abstract or the specific. As we experience a pattern more, it becomes more accessible and more foundational to creating new patterns. In a world where we often encounter combinations of stimuli that seem completely unrelated, these models help us come to reasonable conclusions about our surroundings. If you hear an animal roar but look around and see that you are still in your kitchen, you don’t assume that a bear got into the house, but that somebody has the television on too loud.

The power of these models is apparent: if we know what follows from what we are experiencing right now, we know how to respond. We see a bottle teetering on a table and we can steady it before it falls. We see an old wounds threatening to reopen during an argument and we preemptively make peace. We build tools to initiate a chain of events that lead to a desired goal, whether that is an improved harvest or a man setting foot on the moon. By predicting how the environment will respond to our actions or what state naturally follows from the current one, we exhibit control over reality in a way that simple organisms simply cannot. We nudge the trajectory of the world to be more favorable for us.

What is even more remarkable is that we can create invariant “memories” about things we have never witnessed. We have developed tools, notably language, that allow us to pass on our experiences to others, providing them the ability to respond appropriately to situations before they have ever encountered them. Further, for lack of a better term, we have the ability to imagine. We can create mental worlds rooted in the model of reality that we have built but distinctly different. We can tweak parameters, ask ourselves how things would change if certain patterns coincided in a new way. We can test the outcome of a course of action without taking on the risk ourselves. This gives us our unique capacity for invention in both the practical and artistic senses.

All of this is well and good, and we can provide ample anecdotal evidence to convince ourselves of it based purely on reason. The question is, can it be supported by biology? Unfortunately, there is still so much about the brain, and especially the cortex, that is a mystery to us that little can be said conclusively. Hawkins does, however, offer up a description of what is known about the cortex and how this could endorse the Memory Prediction Framework. This is where my expertise wanes, but I will relay Hawkins’ lesson to the best of my ability:

Physically, the cortex is a thin coating of brain matter, consisting of “grey matter” (neurons) and “white matter” (axons connecting the neurons), surrounding the evolutionarily old brain. The cortex is divided physically into six layers of neurons. Each neuron is connected at many points to the neighbors in its own vertical column as well as to neurons in its neighboring columns and a plethora of other neurons distributed throughout the whole of the rest of the cortex. It is well known that, in the visual regions of the cortex, these connections form further topological hierarchies, and Hawkins thinks it natural that this should also be the case for other regions, as we will see. Signals begin in the lower levels of these hierarchies and then collect and travel upwards as more elaborate relationships between signals are processed.

These upwards connections have been studied widely, but the first important thing that Hawkins focuses on is the fact that there are actually more feedback connections traveling down the hierarchies than there are transmitting data forward. Most theories of the brain seem to discount the importance of these connections, but we will see shortly that they take on prominence in the Memory Prediction Framework.

There is one other important feature of cortical design that we have to discuss first. I already mentioned that the cortex is consistent in its physical makeup across its entirety. In the late 1970’s, this led researcher Vernon Mountcastle to put forth a theory that has largely been dismissed ever since: that there is no significant functional differentiation in the cortex. That is, all regions of the cortex, whether they process sight or language or movement follow some universal algorithm. This, in part, is why Hawkins assumes logical hierarchies within all regions of the cortex and not just the regions that process vision.

As evidence for this claim, Hawkins offers up two arguments. First, the extreme plasticity of the cortex. We know, for example, that violinists have larger areas of their cortex dedicated to controlling the movement of the fingers on their left hand. Individuals born deficient in one sense seem to have larger regions devoted to their others. In one particularly shocking experiment, a man turned blind was able to regain “sight” be having a camera send electrical impulses to his tongue. This somatosensory information was processed in the visual regions of the cortex.

Second, there is no reason to believe that the brain has any reason to process different senses differently. The brain itself is without any sense, after all, and all senses can ultimately be described in the same way. Whether it is a collection of light rays collecting on the retina or longitudinal wave colliding with the cochlea, all of our perception boils down to spatio-temporal patterns. If the brain is just processing spatio-temporal patterns at every turn (as the Memory Prediction Framework already suggests it does), then it should not matter what the input mechanism for these patterns is, just that the pattern is delivered for interpretation.

From here out, I will adopt this theory that all inputs to the cortex are equivalent. I may use language that seems pertinent to the way we perceive a particular sense, notably sight, but please understand that I am referring to any generic sense.

Most of what researchers have observed in the brain has been how these signals are propagated upward through the (logical) hierarchy. A particular image or impression is received from the sensory organs and is transmitted to the cortex. A subset of neurons immediately connected to that sense fires in accordance with the received image and, in firing, pass a signal upwards. The next layer receives this input pattern exactly as the lower layer did and again fires off a subset of its member neurons. This continues until the signal reaches layers high enough for us to make conscious sense of what is being perceived. These groupings of firing neurons, notably the particular groupings that result in recognition, embody the patterns that constitute our invariant memories.

What is interesting is that, while the lowest regions are rapidly modulating due to constantly changing input (because of the saccadic motion of the eye or the continuous flow of sound through the air or whatever else), higher regions -- regions where things in the world are recognized -- stay active for much longer. Obviously, then, these higher regions respond to increasingly general patterns. Hawkins also suggests that there are specific temporal patterns, that is, sequences of spatial patterns, to which these regions respond, and that they will remain active for as long as they receive the same expected repetition of spatial patterns.

Tying into the importance of these temporal patterns, Hawkins now comes back to the significance of the feedback connections, the ones connecting logically higher layers in the hierarchy to those closer to the actual perceived input. In short, Hawkins claim is that, once a pattern has been recognized at a higher level, it tells the levels below which input it is expecting to receive next. In effect, we prime ourselves to see what follows from what we are currently seeing, and we begin responding to it before it can even occur. This is the “prediction” element of the Memory Prediction Framework.

When the actual input defies the expected input, we are jarred out of the current pattern and the image once again propagates normally until we can replace our previous pattern one with one that more accurately represents what we are truly seeing and resume along a more appropriate course of action.

All of this feedback and neural priming predisposes us to see patterns with which we are already familiar. This can help us resolve ambiguities in the environment and perform on-the-fly error correction, but it can also cause us to gloss over seemingly unimportant distinctions. We see this any time we automatically correct a minor spelling or grammatical mistake or even when we try our hands at a “spot the differences” exercise. We are very good at seeing what we want to see because that makes our lives easier and our processing swifter.

So, we now understand how memories are activated and how they are used to predict the immediate future. But how are they formed? Hawkins endorses a simple mechanism known as Hebbian Learning. Simply put, when neurons fire at the same time, the synapses between them are strengthened. This means that each member of a neuron pattern firing increases the odds of the other neurons in the pattern firing as well. These patterns identify an element of a “memory”, and the patterns themselves are stored in the synapses. The more we see a pattern, the more likely we are to see it in the future, even if not all of the elements of that pattern are available in the immediately presented image.

Hawkins has a lot more to say about how the brain operates: the significance of the columnar alignment of neurons as a processing unit; the nature and method of inhibition between neurons and columns; role of the hippocampus as the topmost hierarchical element of the cortex; the importance of the thalamus as a gateway between cortical regions. This is where nuance that is beyond me comes into play, though, and I cannot hope to do all of it justice. I take Hawkins at his word in these arguments, in part because, he offers a number of testable hypotheses that need to be confirmed in order for the Memory Prediction Framework to stand up.

For instance, Hawkins suggests that we will eventually be able to identify the downwards cascades of “predictive” activity through neural hierarchies that should coincide with sudden understanding. Novel events, on the other hand, should be seen propagating upward toward the hippocampus. He identifies specific sorts of cells that should exist in particular cortical layers and which should show excitement in anticipation of an input or respond differentially depending on whether its input is expected or unexpected.

Either because of a lack of interest or continuing limitations in our monitoring technology, it would seem that none of Hawkins’ hypotheses have been either confirmed or refuted in the years since he wrote On Intelligence. If Hawkins really wants to promote his theories or provide some evidence of their correctness without waiting for technology to catch up, the best route might be through implementing them successfully in technology. To this end, Hawkins has already started a company called Numenta, with the goal of producing machine learning packages that model a variation on what he calls Hierarchical Temporal Memory -- a learning architecture that mimics his theories about the design and behavior of the cortex.

Hawkins talks about future generations of machines that utilize Hierarchical Temporal Memory to process patterns and make predictions about any phenomena imaginable. Computers are already better than we are at processing large amounts of data. The electrical signals that travel through a microprocessor are orders of magnitude faster than the electrochemical impulses that drive neural activity. There is no telling what we could discover with intelligent machines churning away and making “informed” predictions based on all of the information we could feed into them.

This is especially apparent when we consider the sorts of input such machines could process. We have already committed to the notion that senses are arbitrary and interchangeable, so why should machines be limited to making decisions based on our senses? Imagine computers that operate and make extrapolations based on input from novel senses, like sonar or barometric pressure, as easily as we do sight. Such computers would drastically increase our ability to, say, predict weather patterns or plan unmanned spaceflight. And that is just the beginning. Imagine how much more effective the approach could be if we actually had computer architectures that integrate memory and processing the way the brain seems to.

But one important question remains: would such a machine be intelligent? This is where we need to start holding Hawkins accountable for some of his early rhetoric. For all of his grand talk early going about the superiority of the brain, Hawkins still appears to believe that intelligence can be replicated by digital machines. This would not necessarily be human intelligence, complete with all of the intangible qualities that make life fascinating, but at least the pattern-driven, future predicting intelligence that he believes is the key to our success as a species. So, whether he commits to it directly or not, Hawkins ultimately believes after all that the brain is nothing more than an augmented Turing Machine; a Turing Machine optimized for the sorts of feedback driven, memory intensive algorithms he has described, but Turing Machine nonetheless.

Moreover, for all of his complaints about behaviorism, Hawkins is ultimately insisting that behavior is the ultimate benchmark of intelligence. He simply shifts the scale from the macro (measurable action) to the micro (method of signal processing). If prediction is the defining measure of intelligence, does the implementation matter? It seems unreasonable to suggest that it is, in which case Hawkins is not really offering us anything new. Prediction, of a sort, has been the goal of A.I. all along. Big Blue was able to “predict” the correct move at each juncture to beat Kasparov. In the abstract realm, the Chinese Room is able to “predict” an appropriate response to a written inquiry in an unknown language. We already have all kinds of learning systems, ranging from traditional Neural Networks to Bayesian modeling, that seek to make predictions based on previous observations without making any effort at faithfully modeling the cortex. Why are these implementations less capable of intelligence than Hawkins’?

Truly, as regards prediction, Hawkins’ solution is different only in approach. This approach may have advantages, it does not give us intelligence by itself. Since prediction alone would appear to be insufficient (unless we suddenly want to change our minds and ascribe intelligence to existing implementations (actual or theoretical) of A.I.), we need to look for another primary feature of intelligence and ask whether that is fundamental to the Memory Prediction Framework and Hierarchical Temporal Memory.

Off the cuff, I would argue this defining feature of intelligence is willfulness -- the ability to direct (if not select) our thoughts, to forge mental connections of our own volition rather than programmatically. Truth be told, I am not certain that the Hierarchical Temporal Memory can provide that or that the Memory Prediction Framework can explain it. Certainly, Hawkins does. He suggests that consciousness is simply “the feeling of having a [sufficient] cortex”. But, while his description of the brain and the neural connections within the cortex certainly provides mechanisms and conduits through which thought could be consciously directed, the driving force is nowhere to be found.

So, if Hawkins’ framework cannot successfully explain intelligence in individuals, we cannot expect it to imbue intelligence on machines. Hierarchical Temporal Memory does not create any stronger argument for understanding than passing the Turing Test. That does not mean his model is useless. The feedback and priming systems that Hierarchical Temporal Memory contains seem well suited to monitoring and making predictions about real time systems, whether it be vision or other crucial signal monitoring. Hawkins’ work may not be any less artificial than other routes that we have taken, but it is still a potentially beneficial supplement.

For the neuroscientist, Hawkins ideas would seem to posses more significant value. The importance of feedback, pattern recognition, and pre-conscious prediction in cognition make a world of sense at a cursory glance, even if the mechanisms that guide it may not uniform throughout the cortex. Even if we limit the Memory Prediction Framework to being a description of cortical function and not intelligence as a whole, though, Hawkins still runs into hot water. For one thing, he discounts the importance of the old brain in intelligence far more than can be acceptable. More than that, though, there has to be a reason that Mountcastle’s theories of universal cortical function have not gained more traction over the past 30-some years.

I may be willing to accept that uniformity could well be the norm for processing sensory input. Patterns in where we process the senses are easily explained by nerve connections from those senses, and all of the signals ultimately are translated into the same sorts neural firings. How does this explain issues like Brocha’s area, though? Why would virtually all humans process linguistic syntax, something with no direct connection to a single sense or the outside world at all, in the same region of the cortex? There does not seem to be an easy “path of least resistance” explination available here.

Since this, ultimately, is Hawkin’s goal, I have to end by giving him credit. His goal was never to provide a definitive answer to these problems, but a starting point. On Intelligence does provide and intriguing, if likely flawed, account of what intelligence could be. I suspect it will not stand the test of time, but it may well provide a useful stepping stone. By taking the risk of proposing not only a theory but standards by which it can be tested, Hawkins has left future researchers with the ammunition to tear apart his framework and iteratively replace it with one that lies closer to the truth. That risk can only lead us closer to the truth.


Bibliography
Hawkins, Jeff. On Intelligence (New York: St. Martin’s Press, 2004).

Hebb, D.O. The Organization of Behavior (New York: Wiley and Sons, 1949).

Mountcastle, Vernon B. “An Organizing Principle for Cerebral Function: The Unit Model and the Distributed System” in The Mindful Brain (Cambridge, Mass: MIT Press, 1978).

Legacy Content. Numenta. http://www.numenta.com/legacy.php. Accessed March 6, 2012.