Wednesday, June 27, 2007

A perspective on persistence

A wonderful fragment from the poem Seventeen Pebbles, from Jane Hirshfield's 2006 collection After: Poems (p. 61)
For days a fly travelled loudly
from window to window,
until at last it landed on one I could open.
It left without thanks or glancing back,
believing only - quite correctly - in its own persistence.

Thursday, June 14, 2007

The Narrative Fallacy, Data Compression, and Counting Characters

I’m very grateful to Tren Griffin and Pierre-Yves Saintoyant for independently suggesting that I read Nassim Nicholas Taleb’s “The Black Swan: The Impact of the Highly Improbable.” Both must’ve realized how relevant his thinking is to my exploration of hard intangibles. At times it felt as if the book was written with me in mind.

One of the human failings that Taleb warns against – “rails against” might be more accurate – is the narrative fallacy. He argues that our inclination to narrate derives from the constraints on information retrieval (Chapter 6, “The Narrative Fallacy”, p. 68-9). He notes three problems: information is costly to obtain, costly to store, and costly to manipulate and retrieve. He notes, “With so many brain cells – one hundred billion (and counting) – the attic is quite large, so the difficulties probably do not arise from storage-capacity limitations, but may just be indexing problems. . .” He then goes on to argue that narrative is a useful form of information compression.

I’m not sure what Taleb means by “indexing”, but I suspect that the compression is required to extracting meaning, not the raw information. It’s true that stories provide a useful retrieval frame; since there’s only a limited number of them, we can perhaps first remember the base story, and then the variation. However, the long-term storage capacity of the brain seems to be essentially unbounded. What’s limited is our ability to manipulate variables in short-term; according to Halford et al, we can handle only about four concurrent items.

Joseph Cambell reportedly claimed that there were seven basic plots, an idea elaborated by Christopher Booker in The Seven Basic Plots; see here for a summary. The number is pretty arbitrary; Cecil Adams reports on a variety of plot counts, between one and sixty-nine. While the number of “basic plots” is arbitrary, the number of key relationships is probably more constant. I’m going to have to get hold of Booker’s book to check out this hypothesis, but in the meantime, a blog post by JL Lockett about 36 Basic Plots lists the main characters; there are typically three of them, or sometimes four. Now, there are many more characters in most plays and novels – but the number of them interacting at any given time is also around four.

I think one might even be able to separate out the data storage from the relationship storage limits: {stories} x {variations} allows one to remember many more narratives than simple {stories}, but I expect that the number of relationships in a given instance of {stories} x {variations} will be no greater that that in a given story.

More generally: if making meaning is a function of juggling relationships (cf. semiotics), a limit on the number of concurrent relationships our brains can handle represents a limit on our ability to find meaning in the world.

Friday, June 08, 2007

Tweaking the Web Metaphor


Today’s modular Internet needs a metaphor make-over. The “silo” and “layer” frameworks that have guided are no longer adequate. It’s time to reinvent a well-worn metaphor: the Web as a web. [1], [2]

The silo model divided up the communications business by end-user experiences like telephony, cable and broadcast television, assuming that each experience has its own infrastructure. The distinct experiences with unique public policy aspects remain, but the silos are growing together at the infrastructure level since all media are now moved around as TCP/IP packet flows. The layer model reflects this integration of different media all using the same protocols. It’s relevant when one takes an infrastructure perspective, but doesn’t take into account the very real differences between, say, real-time voice chat, blogs, and digital video feeds. One might say that the silo model works best “at the top” and a layer model “at the bottom”; in the middle, it’s a mess.

Time to revive the web metaphor, with a twist. The “web” of the World Wide Web refers to the network of pointers from one web page to another. [3] The nodes are pages, and the connections between them are hyperlinks. The “info-web” model I’m exploring here proposes a different mapping: the connections in the web represent information flow, not hyperlinks, and the nodes where they connect are not individual pages but rather functional categories, like blogs, social networking sites, search portals, and individuals.

Food webs

It’s a web as in an ecosystem food web, where the nodes are species and the links are flows of energy and nutrients. The simplest view is that of a food chain: in a Swedish lake, say, ospreys eat pike, which eat perch, which eat bleak, which eat freshwater shrimp, which eat phytoplankton, which get their energy from the sun via photosynthesis. A chain is a very simple model which shows only a linear path for energy and material transfer. (The Layers model resembles a food chain, where network components at one layer pass down communications traffic to the layer below for processing.)

A food web extends the food chain concept to a complex network of interactions. It takes into account aspects ignored in a chain, such as consumers which eat, and are eaten by, multiple species; parasites, and organisms that decompose others; and very big animals that eat very small ones (e.g. whales and plankton). The nodes in a food web are species, and the links between them represent one organism consuming another. While the nodes are multiply connected, there is some degree of hierarchy, since in an ecosystem there’s always a foundation species that harvests energy directly from non-organic sources, usually a plant using sunlight. Each successive organism is at a higher trophic level; first the phytoplankton, then the shrimp, then the bleak, etc.

Info-webs

In an eco-based web model for the Internet, the species in a bio-web are mapped to functionality modules as described in my earlier post, A modular net. For example, a YouTube video clip plugged into a MySpace page running on a Firefox browser on a Windows PC might correspond to the osprey, fish, shrimp, plankton in the simple example above. In the same way that there might be other predators beyond ospreys feeding on fish, there might be many other plug-ins on the MySpace page for IM, audio, etc.. In a bio-web, a link between species A and B means “A eats B”. In the info-web model of the Internet, a link means “information flows from A to B.” Value is added to information value in the nodes through processing (e.g. playing a video) or combinations (e.g. a mash-up). For example, a movie recommender embedded in Facebook gets its information from a database hosted somewhere else, and integrates into a user’s page. Therefore, information transport is key. One can think of the links as being many-stranded if there are many alternative ways of getting the relevant information across, or single-stranded if there’s only one or two communications options (e.g. for web search one can use Wi-Fi, 3G data, DSL, cable modem etc, but for high def video on demand there’s many fewer choices.)

The analogy between the info-web and the food web diverges when one considers what flows across the links. In the Internet, information flows around the web; in the biological case, it’s energy and nutrients. Information can be created at any stage in an information web and increases with each step, whereas energy is conserved, and available energy decreases as one moves up the trophic levels of an ecological pyramid. There is a sequence of “infotrophic levels” where information value is added at each step. However, since the amount of information grows with each step, the “information pyramid” is therefore inverted relative to the ecological one: it grows wider from the bottom to the top, rather than narrower.

Implications for policy making

The Internet is complex web of interlocking service, and is approaching the richness of simple biological ecosystems. In the same way that humans can’t control ecosystems, regulators cannot understand, let alone supervise, all the detailed interactions of the Internet. One may be able to understand the interactions at a local level, e.g. how IP-based voice communications plug into web services, but the system is too big to wrap one’s head around all the dynamics at the same time. [4] This is why a market-based approach is advisable. Markets are the best available way to optimize social systems by distributing decision making among many participants. Markets aren’t perfect, of course, and there are social imperatives like public safety and justice that need government intervention. The info-web model suggests ways to find leverage points where regulators should focus their attention, and also provides salutary lessons about the limits of the effectiveness of human ecosystem management.

For example, a keystone species is one that has a disproportionate effect on its environment relative to its abundance. Black-tailed prairie dogs are a keystone species of the prairie ecosystem; more than 200 other wildlife species have been observed on or near prairie dog colonies. Such an organism plays a role in its ecosystem that is analogous to the role of a keystone in an arch. An ecosystem may experience a dramatic shift if a keystone species is removed, even though that species was a small part of the ecosystem by measures of biomass or productivity. Regulators could apply leverage on “keystone species” rather than searching for bottlenecks or abuse of market power. This would provide a basis for both supportive and punitive action. At the moment search engines are “keystone species” – they play a vital role not only in connecting consumers with information, but also in generating revenue that feeds many business models. One might say that Google is the phytoplankton of the Internet Ocean, converting the light of user attention into the energy of money. Local Internet Service Providers may also be keystone species. In earlier phase of the net, portals were keystone species. Keystone services provide a point of leverage for regulators; they can wield disproportionate influence by controlling behavior of these services.

The unintended side effects of intervention in ecosystems stand as a warning to regulators to tread carefully. For example, the Christian Science Monitor reported recently on efforts to eradicate buffelgrass from the Sonoran Desert. It was introduced by government officials after the Dust Bowl in an attempt to hold the soil and provide feed for cattle. It’s unfortunately turned out to be an invasive weed that threatens the desert ecology, choking out native plants like the iconic saguaro cactus. Another example of biological control gone wrong is the introduction of the cane toad into Australia in 1935 to control two insect pests of sugar cane: it did not control the insects and the Cane Toad itself became an invasive species. By contrast, the release of myxomatosis in 1950 was successful in controlling feral rabbits in that country.


----- Notes -----


[1] Steven Johnson’s Discover essay, republished as “Why the web is like a rain forest” in The Best of Technology Writing, ed. Brendan Koerner, helped inspire this thinking.

[2] This is a rough first draft of ideas. There are still many gaps and ambiguities. The nature of the nodes is still vague: are they applications/services (LinkedIn is one node, Facebook is another), application categories (all kinds of social networking sites are one node), market segments, or something else? How and where does the end user fit in? How can one use this model to address questions of VOIP regulation, accessibility directives, culture quotas for video, and other hot topics in Internet policy? Much work remains to be done. The representation of transport services as links rather than nodes may change. The different in conservation laws needs to be worked out: sunlight, water and nutrients are limited and conserved in the web, rival resources, whereas information is non-rival and can be produced anywhere. Connections need to be made with prior work on metaphors for communications technologies, e.g. Susan Crawford’s Internet Think, Danny Hillis’s Knowledge Web, and Douglas Kellner’s “Metaphors and New Technologies: A Critical Analysis.”

[3] The word web derives from the Old Norse vefr, which is akin to weave. It thus refers to a fabric, or cloth. In many usages, e.g. food webs, there are assumed to be knots or nodes at the intersection of warp and weft, which occur in nets, but not in fabrics.

[4] This is a link to the Hard Intangibles problem more generally, via the limit (about four) on the number of independent variables that humans can process simultaneously.

Thursday, June 07, 2007

Algo Trading

This week's New Scientist has a good review article on algorithmic trading (Robert Matthews, Gordon Gekko makes way for trading software, 30 May 2007).

Some excerpts:

"Investors have realised that the processing speed and sheer volume of trades a computer can make can help them to outwit the sharpest of dealers. . . . Ten years ago, algo-trading was almost non-existent, but according to a recent report by [Brad] Bailey, now at the Boston-based consulting firm Aite Group, one-third of all trading decisions in US markets are now made by machines. He predicts that by 2010 more than half will be done this way. At Deutsche Bank in London, over 70 per cent of a category of foreign currency trades, called "spot trades", are now carried out without human intervention every day."

"Silicon is taking over from carbon on Wall Street," says Bailey.

"Dave Cliff, a computer scientist at Southampton University and founder of Syritta, a UK-based consultancy firm that develops algo-trading software [has turned to genetic algorithms to manage the large number of parameters that have to be tweaked.] His new system takes an initial set of guesses about the optimal selection of market parameters, tests how well each parameter describes prevailing market conditions, and then "breeds" a new selection from these to arrive at a more effective set. This evolutionary cycle is repeated until optimum values for the parameters are reached which the algo then uses to trade with."

"Human traders can make up for the lack of data with instinct and experience, and hooking human instinct up to computing power is now at the leading edge of algo trading. The result is software that helps the trader come up with ideas for bagging some alpha, and tests those ideas in simulations to see if they'll fly. With so many variables, it's easy to make mistakes, but the computer can spot them before unleashing the algo upon the market."

Monday, May 21, 2007

Talk on Hard Intangibles

My recent lecture on Hard Intangibles at the Annenberg Center for Communication at USC (3 May 2007, abstract) has now been posted in a choice of formats:

Audio: MP3
Video: QuickTime or Windows Media

Tuesday, May 15, 2007

A modular net

I concluded in 2005 that modules were a better way to think about the Internet than layers. A book draft by Peter Cowhey, Jonathan Aronson and John Richards has stimulated me to revisit my thinking, since they point out that modularization is a key characteristic of the new evolving Internet.

The problem with Layers

“Layers” were an analytical response to the breakdown of telecoms silos brought about by convergence (Werbach 2002, Whitt 2004). Since the “vertical” silos were merging, a “horizontal” approach seemed like a more useful abstraction, since it would remain relative constant even while technology and business model miscegenation raged between the silos.

Like any model, though, Layers has its limitations. First, layers are premised on a technological abstraction that was honored more in the breach than the observance: after learning about layers in Networking 101, engineers spend the rest of their career using cross-layer violations to improve system performance

Second, there isn’t a single “up direction” for stacking layers. The metaphor of layered stuff presumes gravity, which provides a direction for stacking. Even assuming one can discern an up direction in the network stack (e.g. by referring to successive encapsulation: the network stack is more like Russian Dolls than Legos), there are other important dimensions orthogonal to the network stack. For example, the supply chain from raw material to final product (from up-stream to down-stream, in the lingo) represents an important dimension. We all instinctively use vertical and horizontal as categories, but metaphors of the network stack and supply chains work at cross purposes: upstream in the network business is at the bottom of the stack, and downstream is at the top.

The upstream/downstream dimension (i.e. suppliers/customers) can be contrasted with entrants/substitutes in Michael Porter’s five forces analysis of industry dynamics. The interactions between these forces, as well as competition within an industry, providers of complementary products, and governments, obviate a simple layers model for an industry in general.

Geography and hierarchy are two more dimensions:

  • geography: local, middle mile, long-distance; also related is IXPs where ISPs exchange traffic
  • hierarchy: local, regional, national and international switches
Geography and hierarchy are closely related. However, since one can provide non-local communication without hierarchy, e.g. by using mesh architectures, they’re not identical.

The silo and layer models are attractive because they’re one-dimensional: a single parameter serves to distinguish between categories. In the silo model, the parameter is end-user service (broadcast TV, telephony, satellite communications, etc.); in the layer model, it’s the degree of network abstraction between the physical transmission of data and the end user experience.

Both these parameters (end user experience, and technical architecture) are relevant, so neither silos nor layers are sufficient on their own. Near to the physical layer, a packet is a packet is a packet, and end-user experiences can be more easily ignored; but from the user perspective, watching video is different from voice communications. Layers are helpful low in the stack, but the particular public interest mandates applied to TV are different from those on telephony, and so service silos can’t be ignored. The argument is at its most complex in the middle, where the two blur into each other.

Modules and Modularity

It’s easy to invoke modules and get knowing nods, particularly among techies, but coming up with a comprehensive defintion is tricky. To start concretely, here are some examples of things I count as modules:

  • a network connection of whatever kind, e.g. Wi-Fi access to a base station, a wired thernet connection, a 3G cellular data service
  • directories of all kinds, from the DNS to sites that organize links to other resources, like alluc.org’s pointers to videos
  • a web browser, particularly one which runs on a variety of operating systems, in which an endless variety of services can be run
  • an IM client, particularly on which runs against a variety of IM back-ends, e.g. Trillian plug-ins like web page hit counters, Netvibes widgets, and MySpace templates local and long-distance phone service provision
  • voice over IP
  • A Google AdSense plug-in on a web page

I don’t count these as modules:

  • Cable or satellite TV service (I buy it in a take-it-or-leave-it lump)
  • plain ol’ telephone service (before the AT&T break-up into local and long-distance, and before modems came along)

I’m unsure about some cases. For example, before the advent of plug-ins and RSS, portals like AOL, Yahoo and MSN allowed some on-site personalization, but were pretty much as you found them. You couldn’t mix in third party components to change them, which disqualifies them from modulehood in my book.

With these as a basis, I’d define a module as an interchangeable part of a larger collection of components that delivers an ICT user experience. “Modularity” is the design philosophy which builds functionality out of partial, separable and substitutable components, the modules. The key attributes of modules are:

  • partial – the module is not sufficient on its own to provide a complete user experience; it’s a sub-set of the entire thing. The end user needs to assemble two or more modules to create the result they seek, like combining a local and long-distance phone service provider in the US
  • separable – a module is self-contained and detachable, e.g. a web hit counter or other web plug-in can be removed without affecting functionality of the rest of the page
  • substitutable – a module can replaced by another, equivalent one from another supplier, like replacing one web browser by another one.

Substitutability requires some public disclosure of the interface between modules. This leads to the Farrell & Weiser (2003) definition: “Modularity means organizing complements (products that work with one another) to interoperate through public, nondiscriminatory, and well-understood interfaces.” Note, though, that substitutability is not a sufficient condition; it presumes that an architecture of separable and partial pieces already exists.

Some user-facing modules have enough heft to qualify as applications, though in a modular world they build on some modules, and host others. MySpace or Netvibes are apps, but they plug into a browser, and their functions are extended by other plug-in modules.

Governance

Silos and layers provide straightforward ways to define markets, which can then be used to decide antitrust questions and figure out which groups of players should bear public interest mandates. They were designed to work this way. Modules don’t have this property.

Antitrust remedies are premised on well-defined markets within which companies compete. It’s hard to use modules to define markets, since players can mix and match the modules they use to offer a service, potentially working around a provider with market power. This is good news, of course; antitrust remedies wouldn’t be needed if it’s impossible to put together a platform that forms a bottleneck in the supply chain in a modular world. However, one should Never Say Never; the debate over the proposed Google/DoubleClick acquisition shows that bottlenecks could arise in the Web 2.0 world, too.

There’s also the question of public interest regulation. Universal Service Fund obligations is a well-worn topic, and won’t go away. However, arguably the harder questions revolve around public safety (CALEA, 911, pedophiles), content (obscenity, cultural protection), and access beyond connectivity (e.g. access for the disabled). Who should be responsible for delivering on these mandates, and how? Both silos or layers made it easy: pick a slice, and impose a mandate on companies in that segment. What does one do when not-quite-equivalent end user experiences can be assembled with widely different sets of modules?

My working hypothesis is that regulation should apply to the capabilities that are exposed, not the means by which they’re delivered (e.g. if something’s functionally equivalent to telephony, 911 applies regardless of how it’s done or by whom). Matters are complicated, though, because the context in which a capability is delivered makes a difference (e.g. a voice chat on X-Box Live during a game is different from voice communications module embedded into an employee’s work desktop).

Modules further complicate matters when an end-user builds up an experience by using modules from different providers. For example, imagine a visually disabled person builds a portal on Netvibes with newsfeeds from various web sites, an IM plug-in from one player, and a voice module from another – who’s responsible for delivering accessibility functions?

Odds 'n' Ends

If all modules are one-way pluggable, that is, they form chains without loops, then one can recover a layered categorization. For example, a CNET news feed plugs into my Netvibes page, which plugs into my browser; but CNET feed doesn’t itself host a Netvibes portal. (I feel in my bones that there must be module loops, but I haven’t come up with any yet.)

Modules relate to the interconnection, defined as connections between networks. Interconnect requires “horizontal pluggability” between modules, that is, pluggability among similar modules. The various network transport providers are at the same level of the network stack (i.e. the in- and out-connections use the same protocols) but may be at different geographical and hierarchical levels (e.g. local and long-distance). By contrast, the plug-ins for competing RSS viewers like Netvibes and Google Reader are neither interchangeable across platforms, nor do they directly connect to each other.

Lecture on Hard Intangibles

The Annenberg Center for Communication at USC has posted my lecture on Hard Intangibles (abstract), presented on 3 May 2007, in a choice of formats:

Audio: MP3
Video: QuickTime or Windows Media

Sunday, May 13, 2007

The Perils of Plumbing Parables

No, this story is not about Senator Stevens’ tubes. Jonah Lehrer in The Frontal Cortex links to a Slate story by Darshak Sanghavi on the perils of pluming analogies when thinking about heart attacks.

“It turns out there's a right and wrong place for the plumbing analogy. It's right for people who have heart attacks that involve a sudden, total blockage of a coronary artery. That's why procedures to unclog arteries with expandable stents and balloons ("angioplasty") save lives in emergencies and need to be used more in that setting. But the plumbing analogy fails when applied to stable, partial blockages that don't lead to sudden heart attacks. And yet doctors can't let go of the plumbing talk, and they keep unclogging partial blockages. That's why the vast majority of angioplasties are done for the wrong reasons—that is, for prevention, not acute treatment.”

Ironically, Sanghavi quotes one of his sources invoking a metaphor to explain why preventative angioplasties don’t work: “The trigger isn't bad plumbing—but something more akin to a land mine. People at risk of heart attacks have largely invisible cholesterol plaques throughout their arteries, which act, he says, like unpredictable "little bombs that blow up suddenly and cause a sudden and devastating blockage" in previously healthy-appearing areas.”

More proof, if it were needed, that both lay people and professional decision makers use metaphors to make sense of complex topics, and that models can lead to bad decisions.

Saturday, May 12, 2007

Fingercerting: an alternative to DRM or collective licensing

There was good news for Audible Magic yesterday when MySpace announced that it would use their software to filter out uploads that infringed copyright. Recognizing media clips using fingerprinting (more) has become fashionable as content owners begin to sue hosters.

Fingerprinting, combined with digital certificates, offers a way around the drawbacks of two currently favored ways to govern digital media use. I will focus here on video, since I recently attended a workshop on the future of video copyright at the USC Annenberg Center.

The core of the digital copyright problem is reconciling two valid interests:

Interest 1: Creators’ need to be compensated in order to cover costs and encourage more creation.

Interest 2: Consumers’ ability to make copies of copyrighted material under limited circumstances (loosely, “fair use”).
Here are two fashionable approaches to solving this problem. Each is biased to addressing one of these interests, while ignoring the other.

Solution 1: DRM

In this model, content is locked by DRM under terms specified by the creator/distributor. The consumer can only get access by observing these rules; circumvention is prevented (in the US) by the reverse-engineering terms of the DMCA.

A major difficulty arises in the intersection of DRM with fair use, since the criteria for fair use cannot be encoded in machine-executable form. Thus, Interest 2 above is generally not respected. Other difficulties include the vulnerability to a single hack that puts a piece of content in the clear, particularly if hosted off-shore; the anti-trust consequences of Content/CE/IT standardization; and usability problems with consumer experience.

Solution 2: Collective Licensing

In this model, ISPs would pay a monthly license fee on behalf of each subscriber. This would then be distributed among rights holders a la BMI/ASCAP (cf. EFF’s proposal for music).

A difficulty arises because content creators lose the ability to negotiate their compensation with consumers, thus undermining Interest 1. Owners would be compensated on the basis of some rigid formula determined by the collecting agency. Other difficulties include deriving a formula, since video isn’t as homogeneous as music; anti-trust issues in a collection monopoly; and charging users on enterprise rather than consumer networks. Option 2 also implicitly assumes that DRM is outlawed; if it were allowed to remain, then content creators could get two bites of the apple.

Another way: Fingerprinting + Certificates = FingerCerting

Option 1, the DRM approach, puts the control of content on the user’s device; however, the control is draconian and makes accepted uses like sharing around a user’s personal domain or fair use clumsy at best. Option 2, collective licensing, removes content control by levying a blanket license fee on all broadband subscribers through their ISP, but at the cost of creating an inflexible collecting monopoly and outlawing DRM.

In the “FingerCert” approach, fingerprinting is used to identify content, and an accompanying digital certificate (or “cert”) indicates that the owner has approved its transmission. If the content is registered as copyrighted but not accompanied by a valid digital certificate, an intermediary (ISP or hoster) is obliged to block it. There can still be a negotiation between an owner and a purchaser, but DRM isn’t required, only attaching a cert to indicate a contract. Once the media has been delivered, the cert can evaporate. If the media is provided without encryption, the end user can make copies for fair use without having to worry about arcane and unexpected restrictions.

The big problem with digital media is not personal copies; it’s large-scale illegal distribution. Content owners could use light-weight DRM as a “bump in the road” to mark their rights, but heavyweight (and futile) restrictions intended to prevent even a single hack won’t be necessary. This means a good experience for the vast majority of users who are happy to pay for content, but who would be deterred from buying if DRM were rigorous enough to persuade content executives that their assets were protected against all possible infringement. If you’re only willing to sell sandwiches wrapped in bank vaults, you won’t sell many sandwiches. FingerCerting prevents large-scale distribution by stopping the flow across the Internet, not in someone’s house or between friends’ iPods; it addresses thepiratebay.org and AllofMP3.com, not somebody making a mash-up for their friends. The gates don’t have to be in many places – just the major intersections, like big content sites, or perhaps just at the major IXCs.

FingerCerts gives content owners a way to control distribution of their content (protecting Interest 1), while allowing them to do so without harsh DRM that undermines fair use copying (protecting Interest 2).

What Fingercerting Isn’t

FingerCerting doesn’t require watermarking, that is, embedding (often hiding) a copyright notice in a file. Fingerprinting sets out to recognize the file from its visible characteristics. Watermarking, just like fingerprinting, has to be keep working even when videos are manipulated, e.g. by cropping or transcoding. My uneducated guess is that fingerprinting is more robust in these cases than watermarking since it’s not trying to hide the indicia.

FingerCerting doesn’t require DRM, but neither does it preclude it. It creates an environment where DRM isn’t essential to protecting mass abuse of copyright, and hopefully takes the sting out of the argument over this technology.

Challenges

Any solution to a complex problem will have weaknesses. Here are some I can think of regarding FingerCerting:

Will you need a standard for fingerprints? Audible Magic has a mechanism to register media and recognize clips; so do other companies like Philips. Cert standards exist, but one can imagine different content owners using different solutions. The complexity may be too great for intermediaries if they have to support more than a small number of mechanisms.

Packet inspection technologies to do stream identification are available. Attaching certs to streams is a different issue; I can imagine solutions, but I haven’t stumbled across any yet. Pointers, please.

False negatives – not recognizing an illegal file or stream – will occur, but that’s OK; large scale distribution can stopped since intermediaries will have multiple shots at catching streams. The bigger problem is false positives, that is, when an intermediary mistakenly blocks content. This will annoy users, and present a wonderful scenario for denial of service attacks.

Content hosters/routers will have to be motivated, by litigation or legislation, to implement such a scheme. Current US law provides a disincentive to implementing fingerprinting: Google/YouTube would rather not know that it’s hosting infringing content, because that increases its liability under the DMCA. I presume some legislation or regulation would be required to set up the incentives for a fingercerting process; I don’t know if it will be more or less onerous than that required for DRM (cf. the DMCA) or for collective licensing.

Saturday, April 28, 2007

Mus Gnarus

I made a big deal about the limits to our knowledge of what we don’t know in Incognita Incognita. Reality check: even rats understand the limits of their knowledge, so I shouldn’t get too carried away.

In The Rodent Who Knew Too Much in ScienceNOW (8 Mar 2007, subscription required), Gisela Telis reports on a study that tested the self-knowledge of rats. The experimenters trained rats to understand that they could get a big food reward (six pellets) if they correctly distinguished a long sound from a short one. They could boycott the test if they wanted to, though, and go for a smaller but guaranteed reward (three pellets). As the sounds became harder to distinguish, the rats would opt out and go for the certain, though less generous, reward.

The ability to gauge one’s own knowledge is known as metacognition. We know that humans can do it, and it’s been demonstrated in monkeys and dolphins; this is the first time the effect has been shown in smaller-brained animals.

Metacognition, or “thinking about thinking,” is a strategic activity. It allows us to reflect on a (tactical) cognitive lack, such as having a blind spot for taking immediate steps to remember the name of someone you’re introduced to. (“A pleasure to meet you, John. So tell me, John, did you enjoy the lecture? You know, I agree with that assessment, John.”) Metacognitive strategies can be learned – which gives me hope that any conclusions I might draw about Hard Intangibles will lead to more effective thinking.

P.S. Latin doesn’t seem to distinguish between "rat" and "mouse" (mus). Gnarus means "knowing" or "expert."

Friday, April 27, 2007

Defining the Internet

A discussion started on slashdot last night about a succinct layman’s definition of the Internet.

Most of the examples were technical, and related to computers connected in some way. There were some metaphors: highways, trains, telephones, mail. There were a few references its social aspects. Anthropomorphism was pervasive: computers talking to each other, sending messages, sharing information with each other.

Most of the discussion was about what it was, with some comments about how it works, and occasional references to its social function.

Here’s a précis of the definitions given so far:
  • collection of ideas
  • bunch of connected computers
  • information as trains running on tracks
  • general purpose communication system
  • means for computers to connect to each other and share information
  • everybody already knows what it is
  • computers talking to other computers over cables
  • roadway, highway
  • telephone system with computers calling computers
  • global public computer network
  • mail system
  • physical: computers sending messages; social: virtual community; functional: way to use computers to send messages; technical: computers using protocols
  • agreement (protocol) about how to have networks talk to each other
The best paper I’ve seen on this topic is Susan Crawford’s “Internet Think,” which contrasts the very different ways that “Engineers,” “Telcos,” and “Netheads” define the Internet. Most of the slashdot discussion would fall in the Engineers category.

The metaphors used for the Internet are so stable they’re stale . . . Perhaps the clean slate movement will stir up our thinking. How about the Internet as a brain (back to the Fifties!), or an ecosystem, or a society? The asymmetry of conceptual metaphors is perceptible in the last one: it’s more common to think about society using the Internet as a model (cf. Manuel Castells) than to model the Internet by thinking about society.

Thursday, April 26, 2007

Ducking hard questions: Objective vs. Subjective

In the conclusions to his 1974 paper “Structured Programming with go to Statements,” (Computing Surveys, Vol. 6, No. 4) Donald Knuth observes:

One thing we haven’t spelled out clearly, however, is what makes some go to’s bad and others acceptable. The reason is that we’ve really been directing our attention to the wrong issue, to the objective question of go to elimination instead of the important subjective question of program structure. In the words of John Brown [Knuth citation: “In memoriam . . . .”, unpublished note, January 1974], “The act of focusing our mightiest intellectual resources on the elusive goal of go to-less programs has helped us get our minds off all those really tough and possibly unresolvable problems and issues with which today’s professional programmer would other have to grapple.”

This is a useful and concrete reminder that a fixating on objective, answerable questions can miss the point. There is a certain delight in framing an objective question: it’s elegant, precise, and one can tell when it’s been answered. Some truly important questions, though, don’t lend themselves to objective formulations. This may be because they pertain to complex concepts which have so many interlocking variables that they be reduced to an intelligible logical form, and/or because they refer to notions that are ambiguous or contested. (I suspect that these two conditions, non-linear complexity and ambiguity, are related through our inability to fit the whole of a big question into a single brain.)

Monday, April 16, 2007

Incognita Incognita

It’s hard to think about what we don’t know. If we don’t know something, there’s no “thing” for our consciousness to attend to. One can imagine the unknown as the inverse of what one does know, but that’s just the known combined with the “not” operator, rather than the unknown itself. Most commonly, we tame the unknown with a name. The old mapmakers marked mysterious places as terra incognita, today’s cosmologists explain unexpected galactic dynamics by invoking dark matter, and the religious use the word God.

And of course there’s Donald Rumsfeld, he of the unknown unknown. I’m thinking here of a third category beyond his “known unknown” and “unknown unknown”: the unknowable unknown.

Even though it’s easy enough to think about not knowing, as I’m doing now, it’s not something I do very often. I seldom look at the wall of a lecture theater and realize that I don’t know what’s behind it. My thinking stops at the wall, and bounces back into the room that I can perceive.

Dogs are largely oblivious to human conversation. They don’t follow the to and fro of conversation. They are aware of the sound and some if its import, but they don’t know its meaning. In a sense, it doesn’t exist for them. In the same way, I’m ignorant of much going on inside me and around me, and I’m ignorant of the fact that I’m ignorant.

Things I know sometimes feel like places. As I learn more about a subject, I can begin to assemble the rooms representing topics into a building. But if I don’t know something (statistics, say) it’s not as if it’s the unexplored south wing of a mansion. There is no south wing. I have no sense of its shape. Something once known but now forgotten (like Green functions, in my case) are ghostly ruins remembered from a dream; there are only wisps and fragments.

We make up stuff to hide the fact that we don’t know. Helen Phillips describes in New Scientist (“Mind fiction: Why your brain tells tall tales,” 7 October 2006) how people make up stories when the reasons for their action are not available to conscious introspection:
[Timothy Wilson and Richard Nisbett] laid out a display of four identical items of clothing and asked people to pick which they thought was the best quality. It is known that people tend to subconsciously prefer the rightmost object in a sequence if given no other choice criteria, and sure enough about four out of five participants did favour the garment on the right. Yet when asked why they made the choice they did, nobody gave position as a reason. It was always about the fineness of the weave, richer colour or superior texture. This suggests that while we may make our decisions subconsciously, we rationalise them in our consciousness, and the way we do so may be pure fiction, or confabulation.

Note that people didn’t say, “I don’t know.” This is an important result for the study of hard intangibles. We are usually not aware that we have a limitation. Sometimes cannot even believe that we’re limited. Here’s another excerpt from the New Scientist story:
It is surprisingly common for stroke patients with paralysed limbs or even blindness to deny they have anything wrong with them, even if only for a couple of days after the event. They often make up elaborate tales to explain away their problems. One of Hirstein's patients, for example, had a paralysed arm, but believed it was normal, telling him that the dead arm lying in the bed beside her was not in fact her own. When he pointed out her wedding ring, she said with horror that someone had taken it. When asked to prove her arm was fine, by moving it, she made up an excuse about her arthritis being painful. It seems amazing that she could believe such an impossible story. Yet when Vilayanur Ramachandran of the University of California, San Diego, offered cash to patients with this kind of delusion, promising higher rewards for tasks they couldn't possibly do - such as clapping or changing a light bulb - and lower rewards for tasks they ould, they would always attempt the high pay-off task, as if they genuinely had no idea they would fail.

If we can observe the limitation in others, we can at least study it, if not experience it ourselves. However, it will be hard to teach others – and ourselves – to behave differently if, in our bones, we still cannot conceive of our lack.

Tuesday, April 10, 2007

Let’s hope we’re not rational about climate change

Global warming is a classic collective action dilemma.

A solution to global warming is a collective good and will be undersupplied, as Mancur Olson pointed out back in 1965.

Therefore, if Olson’s premises and argument are valid, we’re dooooooomed.

However: his argument supposed a rational economic agent who will wait for others to act, since his contribution is so small that on its own it won’t make a difference, and it’s absence won’t be noticed.

Only if humans don’t act as selfish rational agents will we avoid a climate catastrophe.

Fortunately, behavioral economics etc. suggests that we have bounded rationality, and even better, psychology and evolutionary biology suggests that non-rational altruism is hard wired.

Maybe there’s hope.

Tuesday, April 03, 2007

Algorithmic trading changes markets (maybe)

Kyril Faenov alerted me that electronic exchanges feeds back on themselves in unprecedented ways due to automated (or algorithmic) trading. Traders observe the market, and imagine a way to make money; their quants then write software to execute this trading strategy automatically. Running this code creates new, fast and extensive linkages between market processes.

Robin Sharpe provides an excellent introduction to algorithmic trading in Automated Trading and the New Markets. For more information, see John Bates in Dr. Dobbs, and wikipedia.

For example, trader A (or their software) notices a periodic spike in the price of equity X; trader B may be trying to buy a very large position of X in small portions so as not to push up the price too much. Trader A (or the software) buys stock X just moments before each predicted spike, selling it to trader B at a higher price once B enters the market. Trader B (or their software) notices the run-up, and changes their buying rhythm to disrupt trader A. And on it goes. . .

Such market interactions aren’t new, but software can execute the trades faster than humans can respond to them. Rather than duels between traders in real time, it becomes a duel between traders’ models of how the market functions. These are contests between world theories, where the theories themselves constitute the world. Kyril calls this the “reflectivity” of the market.

Douglas Hofstadter introduced the term “strange loop” in Gödel, Escher, Bach to describe a series of steps through a hierarchical system which take one back to the beginning. His new book I Am a Strange Loop uses this concept to explain self-awareness. In a New Scientist interview, Hofstadter says the brain, and the self, is like a smile because it’s a process rather than a thing. (Extending the metaphor: Software is to hardware as a smile is to a body.)

A market is observable through its behavior, that is, how it responds to stimuli. When the responses happen faster than humans can follow, and involve the integration of more variables than humans can handle on their own, the market is less a social interaction among people than an environment in which people act. A market is neither a place, nor a group of traders, nor the sequence of trades, nor a reflection of an outside commercial reality; it is the self-perpetuating process that involves all of them. The system’s behavior becomes a subconscious expression of the cumulative conscious plans of many people – subconscious because the mechanism is not directly available to human introspection.

Markets are examples of distributed cognition, that is, cognition which occurs in an ensemble of people and tools, rather than in a single brain; see Giere (2002, PDF) for a good survey. What’s striking about algorithmic trading is that the amount of cognition occurring outside human brains is growing rapidly.

Robin Sharp (ibid) points out that trading is increasingly hands-off, since humans can’t cope with the reaction times required.

“Whilst the theory behind program trading is fairly simple, the software reality is that program trading operates at a different time-scale to even the fastest human trading. . . It’s worth understanding that the brain brings experience and subjectivity to the table and software brings speed and objectivity to the table. For most types of trading experience beats speed but there is a lot of noise in the market in the sub-second region where the brain simply can’t compete with a computer. . . At the moment the volume of trading at these sub-second time scales is not be great (less than 5%) and is held back because the coarse granularity of ticks and price data has been designed for human interaction. However this will change.”
He also notes that algorithms can exploit the capacity limitations of human traders: “A large number of trades are difficult for traders to juggle in their heads. When such large events occur is small time frames computers can predict irrational behaviour of traders, again for a profit.”

The kinds of problems that traders face are not only analytical or cognitive; they’re also social. Sharp notes the organizational impediments to certain kinds of trades: “One of the reasons traders don’t trade cross market is that they cannot price the instruments fast enough, and – again – is that traditional corporate management structures and regulatory structures impede cross-market trading.”

Sharp tabulates the latencies in a trading cycle. The computer processing is 200 milliseconds, quick compared to 1,250 milliseconds for human perception, evaluation and response. The network latency is about 400 milliseconds. It’s worth noting that these network applications reinforce old geographic patterns rather than abolishing distance. Kyril tells me that the NYSE is doing a good business selling rack space at the Exchange to big trading houses; because every millisecond counts, their computers need to be on the same LAN as the Exchange or else they’ll lose to faster arbitrageurs.

I’m struggling with the question of whether automated trading leads to a qualitative change in the markets, rather than simply a quantitative one. While algorithms that respond to changes in the market begin to constitute the market, this kind of loop applies to traditional human-only trading, too.

Sure, the feedback loop is much faster with automated trading. But is it a difference in degree, or a difference in kind? It’s different from the human perspective; stuff now happens too quickly to follow consciously, and the role of humans changes. However, it might simply be a change in time scale, not a change in process.

The increasing complexity of the market may be even more important. Arbitrage can link markets, which generates more correlated variables than traders can juggle in their heads. An automated trading strategy could link current and futures markets, on different exchanges (New York and London), for different instruments (equities and foreign exchange), and different data types (Reuters news feed, GOOG and MSFT stock prices, S&P500 index, the 15 minute volume weighted price of GOOG). As Robin Sharp points out: “Eventually arbitrage will force separate markets to revaluate their relationships. It only takes one successful arbitrage engine to forever link two previously unrelated markets. Anybody in the business will know how deeply this will be felt.”

Perhaps integrating more streams of information in more complex and rapid ways creates a new kind of market causality. When humans were trading with each other, it was a social process. Now, from the human perspective, it’s more like experimenting on the world than dealing with people. “Hard intangibles” come into play because this new world is not the one humans evolved in. Humans still set the goals and strategies, but the parameters of this world interact in unexpected ways. And genetic algorithms will lead to algorithmic trades that are profoundly alien to human intuition. Trading is another activity, like large software projects, where the abstractions we’ve created are beginning to outstrip our ability to understand them.

------------

Ronald Giere, “Scientific cognition as distributed cognition,” in The Cognitive Basis of Science, Cunningham, Stich & Siegal (eds.), Cambridge University Press, 2002