Showing posts with label Big Data. Show all posts
Showing posts with label Big Data. Show all posts

Sunday, June 9, 2013

Panopticon - American Style

The myriad (and growing) surveillance leaks of the past few weeks put me in mind of Foucault's thoughts on panoptic surveillance.

  The idea of the Panopticon traces back to the work of Jeremy Bentham in the 18th century.  Bentham coined the term to describe a design for prisons where every inch of every cell was under constant surveillance by prison authorities; yet prisoners would not know when, or if, they were actually being watched.  Bentham held that unverifiable surveillance was an efficient and economic exercise in state power, while providing the state with a viable mechanism for discipline and control.

French philosopher Michel Foucault used the idea as a metaphor for the modern state's mechanism for control and punishment.  Under Benthan's unverifiable surveillance, individuals are never sure whether or not they are being watched, or by whom.  As a result, Foucault argues that individuals are "trained" to resist any impulse of misbehavior or "abnormality" for fear of being caught, uncertain whether such a display is safe or not.  Also, as more and more information is gathered on individuals and their behaviors, Foucault argued, people become "objects of knowledge" to state authorities - whose actions can be tracked, examined, catalogued, examined, compiled, and acted on - if that is in the interest of the state.  Foucault argues that this gives hegemonic authorities great power - to identify, track, and punish those whose behaviors lie outside the state's idea of "normal" - and allowing the state to "humanely rehabilitate" those of their citizens who stray from the path of enlightenment (or at least what the hegemonic authorities think is appropriate).

I had originally put Foucault's social theories among those critical theorists who had interesting ideas and insights, but whose insistence on the prevalence of a powerful hegemonic state authority seemed unrealistic.  Then came the leaks and the leaks about the state's response to leaks.

First there were the revelations of the Department of Justice intercepting and tracking phone calls and emails of journalists who broke and wrote about various "leaks" that weren't seen as beneficial to the current administration.  Then came the Guardian's story about the state's downloading of all call information from one of the US's largest operators - Verizon.  The FISA warrant was notable for its breadth - instead of looking for individual records, or those meeting certain behavior patterns or posing identifiable threats, the NSA sought, and obtained, permission to gather all records for everybody - who called whom, where calls originated and terminated, and length of the call.  Also included were customer records, emails, and Internet activity routed through Verizon's system.  Further, all of this was supposed to occur without any notice being given to the individuals being surveilled.  A perfect example of the concept of panoptic "unverifiable surveillance."  [For those of you on other networks, Verizon's just the one the published the warrant for - it's a safe bet that similar warrants have been granted for most, if not all, telecommunications operators in the US.]

If that level of unverifiable surveillance wasn't bad enough, there was PRISM.  The Prism project actually consists of a number of different efforts with a common goal - to gather, copy, and analyze all electronic communication.  There's a physical basis in the form of a large complex for data storage and analysis, and it's known that other aspects involve collecting the information running through various checkpoints on the Internet, ostensibly from, and with the cooperation of, major Internet players (who have subsequently vociferously denied any cooperation with, or even awareness of, the program).  The Guardian reported that one of these programs, called "Boundless Informant," collected more than 3 billion pieces of information from U.S. computer networks last March (2013), and that the intercepts do include specific information such as IP addresses.  All for a program that is supposed to be limited to foreign communications.
  Whether it is focussed terror-related intelligence or a broad and comprehensive surveillance of all electronic communications - including those of its citizens remains a bit unclear.  But NSA testimony to Congress has shifted - from insisting that it does not collect any type of data on large numbers of American citizens, to that the NSA had not "wittingly" (purposely) collected data on large numbers of Americans, to tightly parsed denials that the NSA could, in fact, "physically" collect very specific bits of information (the kinds of things that you can't tell from online communications).  And in a classified slideshow the Prism folks use to market their activities to other agencies, they clearly suggest that they can deal with most all forms of electronic communications.

Welcome to the world of unverifiable surveillance.  But the Foucault conception of panoptic discipline also requires the ability to take all that surveillance data and use it to identify outliers - those who stray from social normality.  And that brings us to the data analysis side of Prism, and the developing field of data mining.  One of the things that data mining is good at is identifying outliers - individual cases that vary significantly from the vast muddle in the middle that we can consider typical or normal.  In other words, we're becoming very, very good at treating people as "objects of knowledge", and being able to recognize and identify outliers.

Finally, the Obama administration has amply demonstrated its interest in disciplining those who stray from the enlightened path, and not necessarily through humane rehabilitation.

Maybe Foucault's fears are more realistic and immediate than I originally thought.

Tuesday, March 19, 2013

The Internet, Big Data, and the Surveillance State

In a recent opinion piece for CNN, Bruce Schneier proclaimed: "The Internet is a surveillance state."

The Internet's never been secure or private - by design.  It's design goals were to be open and shared - to make it universally accessible, flexible, and adaptable.  And for those who remember the DARPA (defense-related) roots, even there the primary goal was survivability rather than security.  There's a reason the military's never relied on it.
  Sure, there are things you can do to make Internet use somewhat more private - use encryption, route through anonymizers, etc.  But still, every bit of data carries addresses, and all that flexibility and sharing requires that basic information on users and connected devices be readily available.  Add the fact that every data packet travels public routes where they can be duplicated, and ISPs and servers regularly back-up content and messages, and you realize that the Internet is a very public place.  As for encryption, industries trying to rely on encryption for copyright protection (as well as governments) have found that every encryption system is beatable, given enough brains, computing power, and time.  That many governments seek to restrict the use of encryption technology is a matter of laziness and cost rather than a fear of totally private communications.
  For a long time, the sheer volume of Internet traffic provided a bit of privacy protection for common users - searching through the volume of packets and files, identifying and matching traffic through multiple sites, etc. was just too problematic.  But if you had the resources, you could often break through whatever privacy/security roadblocks used (if any).  Schneier offers three recent illustrations -
  • the Chinese military hackers that have been attacking U.S. and European government, military, and commercial sites, were identified in part as they accessed their Facebook accounts through the same networks and hardware used for the hacking.
  • a leader of the LulzSec hacker collective was identified and arrested, reportedly because he slipped up and once logged into an IR chatroom in the clear - without masking his IP address as was his normal practice.
  • Paula Broadwell, who had an affair with then CIA Director David Petraeus, was identified despite only logging into the anonymous email account created and used for the affair from public internet sites.  The FBI reportedly identified her by matching hotel and service receipt records from the times of the emails, and finding hers was the one name in common.
Schneier's point is that Internet traffic is widely tracked, and not only by governments and counter-espionage organizations.  Google does it on everything running through one or another of their sites.  Google also tracks and records websites and content for its search engines.  Blogger.com, for example, lets me know who's visited this blog, where you're from, what OS you're running, and how you found me.  Apple tracks user behaviors on iPhones and iPads.  Facebook tracks its members and their behaviors, and backs up their content and submissions.  They've also admitted tracking their members non-Facebook activities, and using cookies to track online behaviors of non-members who visit Facebook pages.  Pretty much every commercial site builds profiles of users and customers.  And metrics firms collect and track data (anonymized, they say) on users and their Internet behaviors in terms of data traffic flows..

Now if all these were separate, private, and secure, they may be seen by many as the acceptable cost for the services and benefits provided by the Internet and various online services.  Even if they were shared, it might not be so bad, if it would take significant time and effort to try to link things together (particularly if you're looking for patterns in behavior).  If "surveillance" was too costly or inconvenient to be used regularly or for trivial purposes.
However, that's increasingly not the case, due to technology advances and the rise of Big Data.  If you haven't heard the phrase before, Big Data refers to a range of programs and techniques for trolling extremely large data sets (such as online tracking data) to tease out and identify patterns and links.  With Big Data to help, the sheer volume of online data is no hindrance.  Automated systems can scan millions of emails in real time looking for key words or phrases.  Automated systems can match online searches, or the use of certain apps, to purchasing behaviors and location data from mobile devices to send users a coupon for a nearby store or restaurant.  And data storage costs keep falling.  (And while not exclusively Internet, facial recognition software and the myriad private and public video cameras can be used to track a person's movements).
  As Schneier puts it,
This is ubiquitous surveillance: All of us being watched, all the time, and that data being stored forever. This is what a surveillance state looks like, and it's efficient beyond the wildest dreams of George Orwell.
Nor does there seem to be an easy solution, or a means of opting out.
There are simply too many ways to be tracked. The Internet, e-mail, cell phones, web browsers, social networking sites, search engines: these have become necessities, and it's fanciful to expect people to simply refuse to use them just because they don't like the spying, especially since the full extent of such spying is deliberately hidden from us and there are few alternatives being marketed by companies that don't spy.
So, Schneier concludes, welcome to an Internet without privacy; welcome to the Internet surveillance state.  While public interest groups try to raise concerns about privacy, and individuals rant, the public doesn't seem to mind - as long as Amazon and Netflix make good recommendations, YouTube lets you know about the latest "cute kitty" viral video, and social media don't charge fees.  The Internet was never truly private in the first place, and isn't likely to ever significantly shift in that direction.  In part because one of the significant public values of the Internet comes from the lack of privacy and the ability to find and make connections.  What international and national regulatory moves there are are about giving governments more control over the Internet, and more access to the information it transmits and generates.  Which means even fewer real privacy protections.

  If that worries you - and it should - you could go offline.  But in a modern global information society, that comes at a high cost.  Or you could try to level the playing field, as David Brin suggests in The Transparent Society - let us, as citizens, have the same access to surveillance of government activities as the government has over our activities.  Make government truly transparent, rather than settling for "transparency" being defined as giving people access to information the government wants to provide them.  Turn the cameras around, open records, and let the public see what government actually does, rather than only what the government claims it's doing (true or not).  Or hoping that an (increasingly scarce) honest and aggressive press will investigate and report, and do the monitoring for them.

Sources -  The Internet is a surveillance stateCNN Opinion
David Brin's Transparency website

Edit - fixed some typos

Sunday, November 11, 2012

Big Data and Gamification

How can you make sense of Big Data, the reams (or gigabytes) of data that's generated with online activity?  One new approach seeks to make it all a game - of sorts.  Gamification is the term being used for the application of gaming principles to non-game applications - such as Big Data analysis and tracking.
  Badgeville is a new start-up that's marketing gamification apps designed to measure and influence user activity on Web and mobile sites.  For example, Badgeville borrows tools from social games to better understand users and to spur desired user behaviors. For example, gamification techniques could assign points for specific actions, recognize achievements, create user contests to unlock awards, and provide real-time notifications when users perform a designated activity.  CEO Kris Duggan explains:
"A lot of people are talking about big data -- collecting data and profiling users, all these kinds of things. But they don't really know what to do with that data, and they don't know how to make it actionable... We track all users and what they're doing," he explained. "This allows us to understand what behaviors users are performing, and what motivates their behaviors."
  Such things can be very useful for media and information apps, particularly those that seek to use their websites and mobile apps to engage with and build audiences.  Games or game-like apps can be very attractive means of audience engagement.  Getting more info on user attitudes and behaviors will likely be even more important, allowing companies to better target audiences in a crowded and increasingly competitive market.  A number of big media firms (NBC, Washington Post, Fox)are already working with Badgeville - hopefully we'll get some idea of how effective ramification can be for media outlets.

Source  -  Big Data and Gamification - Good Partners?  Information Week

Tuesday, August 14, 2012

Two Posts on Big Data & Social Media

The Internet allows for the collection of immense amounts of information - on users, on outlets, and on content and its flows.  Aside from continuing concerns about privacy issue, there's been a rise in the attempts to make sense of, and use, all this constantly-generated "Big Data."  Furthermore, the rise of social media, and its transparency, is creating huge amounts of comments, thoughts, recommendations, of millions of users - as well as their follies and foibles - all of which can be mined for data.  
  We've clearly past the first stage - the explosion of information.  That's been happening for decades, and the amount of information generated around the world continues to explode exponentially.  We're also well into the second stage - developing tools and techniques for sorting, managing, and even filtering information.  We're also in the early stages of developing analytics - the means to measure and analyze all that information (although measuring compounds the information explosion as it continuously creates new data about information and data).  And that leaves us with the continuing problem of making sense and finding value.  We're making inroads, but have yet to fully step into the third wave - using all those tools to find value in the information haystack.

Dion Hinchcomb, posting at The Brainyard, argues that
the social world, by dint of a billion people engaging with each other around the clock, is now the richest source of open innovation, product ideas, marketing and sales opportunities, customer care capacity, and much more. One thing we've learned in the last eight years of the mass collaboration era is that, whatever an organization cares about, crowds can help us conceive of it, build it, test it, market it, support it, and fix it--and do all of that at scale.
The problem's been to find the gems or spot the trends in this morass of information and data.  Thankfully, there's been a lot of people and companies working to develop analytics and techniques to sift and sort Big Data. This has opened Pandora's Box - a potential of finding value for Big Data users, as well as the potential for harming social media and Internet users.  Hinchcomb focuses on the positive, positing that Social Media's Big Data can generate positive returns on organization's investment in utilizing and analyzing social media.  Hinchcomb suggests that we're well past the first wave - that organizations are finding and making use of social media.
With the continuing rapid growth of social media, Hinchcomb suggests that some organizations will transition to social businesses.  In an information economy, knowledge workers recognize the benefits of the opportunities the Internet and social media provide - greater access to information, enhanced opportunities for collaborations beyond your own "silo" of expertise/focus, and the ability to focus on project-related tasks.  Particularly if enterprises can transcend internal barriers, such as embedded legacy enterprise-specific applications.  Still, the fairly rapid adoption of social media provides hope that we'll make the transition.

Still, it's not all blooming roses out there in the world of Big Data.  At a recent Kontagent Konnect user conference, Josh Williams talked about the "Seven Deadly Sins of Data Science."  As with any analysis, you can do it well, or poorly (particularly if you don't understand the limits inherent in any analytic technique).  Here's some of the possible ways to mess up.
  1. Sloth - Lazy Data Collection:  Also known as GIGO (garbage in, garbage out), the first limit on analysis is the quality of the data.  It's easy to grab and use numbers that are there, rather than the numbers that you need.
  2. Negligence - Misapplied Analysis: It's easy to use an inappropriate technique - one that doesn't provide relevant results. (An example from one of my first stats classes - "The average American is 47% male and three days pregnant."  Think about it.)
  3. Gluttony - Too Many Reports: Too much information and too many tools can result in too much analysis - and the key results can get lost in the mix.
  4. Polemy - Data Definition, User Disagreements:  One of the problems with "Big Data" at this stage of development is that there are few widely-accepted metrics or analytics, which contributes to arguments over the definition and utility of data measures and procedures.
  5. Imprudence - Jumping to Conclusions: When you have a lot of data and a lot of results, it's easy for something to jump out and seem significant.  Stats people will remind you that even at 95% confidence level, there's a 1 in 20 chance of a false positive (i.e., what you see isn't really there).  If you place too much importance on a single finding, without examining the broader context, definitions, or limits of measures, you could be making a big mistake.
  6. Pride - Decision-Driven Data Making:  If you look at Big Data to confirm your beliefs, it's easy to construe or manipulate things, even subconsciously.  You might define things a certain way, pick supportive data, manipulate datasets - all of which might bias the analysis to foster confirmation.  The true scientist looks for answers rather than confirmation, and is more likely to get to the truth (or reality or whatever).
  7. Torpor - Learning and Acting Slowly:  Historically, data and analysis have been delayed - collecting, publishing, and analysing data took time (in academia you can easily have delays of 1-2 years in getting results out and being able to apply them).  However, the Internet, social media, and Big Data are all racing in real-time.  Being successful there requires collecting and examining data in real time, and being able to react to what you see happening quickly.  A delay can put you on the wrong end of the trend.
What all of these suggest is that when seeking to use Big Data or social media, you need to learn the lessons of data definition, data collection, and data analysis that decades of research and practice in other areas have yielded.  But you also have to be aware of the unique aspects of Big Data, social media - particularly the speed at which things occur online.  Being aware of the 7 Deadly Sins and how they might lead to poor results can help.

Sources - Why Big Data Will Deliver ROI for Social BusinessThe Brainyard
The Three Waves of Enterprise 2.0: Climbing the Social Computing Maturity CurveebizQ
Social business holds steady gap behind consumer social mediaZDNet
7 Deadly Sins of Big Data Users,  Information Week

Thursday, January 5, 2012

Big Data and Journalism

With the rise of cheap computing and data storage has come the ability to measure and store huge amounts of data.  What's coming along a bit more slowly is the interest in, and ability to make use of all of that data.  Another jump in the ability to make use of all that info came with distributed computing - first with standalone projects like SETI@home, which used the processing power of millions of home computers to process billions of pieces of data (2 billion so far), and now the ability to harness the thousands of virtual computers in the Cloud.

So what is  Big Data and what does it have to do with the future of journalism?  The "Big Data" concept refers to the tools and processes for managing and using large datasets.  The idea of data-driven journalism, has been around for decades, but for the most part been limited to focused use of datasets to answer specific questions.  And, quite frankly, it's been severely limited by most journalist's seemingly inherent antipathy to numbers and math, as well as the decline in investigative journalism.

More recently, the concept of database journalism has emerged.  Unlike data-driven journalism, the idea of database journalism is to aggregate the materials collected by journalists into databases, which can then be used to spot trends or provide local illustrations for local versions of stories.

Neither of these fit the idea of Big Data, however.  What the Las Vegas Sun is doing with data may qualify, though - they exploit the massive amounts of audience metric generated by their online edition to suggest coverage, link to public databases to generate real-time informational maps of things like police reports, real estate listings, and retail hours  for local editions, and used public and online databases to research a story on local healthcare.

But there is the potential for much more - particularly in today's age of big data and huge document dumps (often designed to hide the big stories from easy access).  Only traditional journalism hasn't had a lot of interest in, or ability to exploit, Big Data.  From various Wikileaks dumps to the release of Stimulus-funded projects data, to Sarah Palin's emails, journalists have let others do the analysis and largely just reported what they were told (if they reported it at all).  That's a shame, because there is an unprecedented amount of publicly available information on government activities at all levels, campaign contributions and links between big money and "independent" public interest groups that should make a "watchdog" press drool.  Not to mention how monitoring search engines and social media could alert journalists to emerging issues and hot topics.  (For example, Google does a faster and better job of tracking flu outbreaks than the CDC, simply by monitoring searches for "flu remedies" and "flue symptoms.").

If journalism is to have a future, they need to do more than simply report what others say and do - they need to originate news, add value to stories, and reveal the needle in the haystack.  And doing that through Big Data, through the use and analysis of available information, is becoming easier and cheaper.  Will journalists acquire the interest and skills to do so, or will they leave that to others? (and in doing so render themselves even more irrelevant).

Source -  Big Data: Why All the Fuss?  InformationWeek Global CIO