Showing posts with label online data. Show all posts
Showing posts with label online data. Show all posts

Monday, December 9, 2013

Viva la Revolution Mobile

From CIO Insight - 10 Awesome Facts About the Mobile Revolution slideshow.

Several highlight the scale and scope of mobile, mobile data, and mobile broadband
  • There are more than 6.7 global mobile subscriptions, 30% for smartphones
  • Global subscriptions for mobile broadband will pass 2 billion this year, projected to hit 8 billion by end of 2019
  • Subscriptions for mobile PCs, tablets, and mobile routers are projected to grow from 300 million currently, to 750 million by 2019
  • In 2009, there was more voice traffic on the mobile network than data traffic.  Today, data traffic is more than 9 times higher than voice traffic.
  • Data traffic per smartphone currently averages 600 MB per month, expected to hit 2200 MB per month by 2019
  • Data traffic per tablet currently averages 1000 MB/mo; will hit 4500 MB/mo  by 2019.  Mobile PC data traffic currently averages 3300 MB/mo, projected to reach 13,000 MB/mo in 2019
  • By 2019, 95% of North American will have LTE (mobile broadband) coverage.  Globally, about 2/3 of the world's population will have LTE access.
A significant contributor to the mobile revolution is the growth of online video options and services, social media, and various entertainment/gaming options
  • By 2019, half of mobile data traffic will derive from video.
  • Smartphone owners spend, on average, 13 hours monthly on social networking, 8 hours on entertainment, and 6 hours gaming.
The impact may be greatest in Asia, where the growth of mobile broadband provides quick and relatively cheap access to the Net and its content and services to some of the world's most populous countries.
  • China alone has 1.2 mobile subscriptions, and India 742 million.  The rest of the Asia-Pacific region includes a further 1.3 billion subscriptions.  In contrast, Africa accounts for 803 million subscriptions, and Latin America 697 million.
The numbers and forecasts come from the most recent Ericsson Mobility Report.

Source -   10 Awesome Facts About the Mobile Revolution,  CIO Insight
Ericsson Mobility Report - November 2013

Tuesday, May 7, 2013

Poynter - Digital Tools for Journalists

Poynter's website has a good post on 10 digital tools they feel can help improve reporting and storytelling.
  Several are what I'd call research tools - FOIA Machine | (@FOIAMachine) (helps create and track FOIA requests); Census.IRE | (@IRE_NICAR) (search, organize and view 2010 Census data); iWitness | (@AdaptivePath) (curates social media, searchable by time and location). 
   However, most are storytelling resources; things that can provide background, graphics, and facilitate multimedia storytelling.  
Public Insight Network | (@publicinsight) and Ushahidi | (@ushahidi) maintain databases of people eager to tell their story - although I'd worry about accuracy and bias in what is essentially advocacy.
Several help journalists create and develop graphics or incorporate multiple media forms into stories - TileMill | (@TileMill) (creates interactive maps); Tableau Public | (@tableau) (user friendly app for compiling graphs, charts, etc.); Popcorn Maker | (@mozilla) (adds interactive features to videos); Atavist | (@theatavist) (tool to compile elements into a single multimedia story, app, magazine, or ebook.
The final one is The PANDA Project | (@pandaproject), which allows newsrooms to share data and materials among reporters, and allows online collaboration.

I'll encourage you to check the Poynter post for more details, and examples of some of the ways media outlets have utilized the tools.

Source -  10 digital tools journalists can use to improve their reporting, storytelling,  Poynter

Tuesday, March 19, 2013

The Internet, Big Data, and the Surveillance State

In a recent opinion piece for CNN, Bruce Schneier proclaimed: "The Internet is a surveillance state."

The Internet's never been secure or private - by design.  It's design goals were to be open and shared - to make it universally accessible, flexible, and adaptable.  And for those who remember the DARPA (defense-related) roots, even there the primary goal was survivability rather than security.  There's a reason the military's never relied on it.
  Sure, there are things you can do to make Internet use somewhat more private - use encryption, route through anonymizers, etc.  But still, every bit of data carries addresses, and all that flexibility and sharing requires that basic information on users and connected devices be readily available.  Add the fact that every data packet travels public routes where they can be duplicated, and ISPs and servers regularly back-up content and messages, and you realize that the Internet is a very public place.  As for encryption, industries trying to rely on encryption for copyright protection (as well as governments) have found that every encryption system is beatable, given enough brains, computing power, and time.  That many governments seek to restrict the use of encryption technology is a matter of laziness and cost rather than a fear of totally private communications.
  For a long time, the sheer volume of Internet traffic provided a bit of privacy protection for common users - searching through the volume of packets and files, identifying and matching traffic through multiple sites, etc. was just too problematic.  But if you had the resources, you could often break through whatever privacy/security roadblocks used (if any).  Schneier offers three recent illustrations -
  • the Chinese military hackers that have been attacking U.S. and European government, military, and commercial sites, were identified in part as they accessed their Facebook accounts through the same networks and hardware used for the hacking.
  • a leader of the LulzSec hacker collective was identified and arrested, reportedly because he slipped up and once logged into an IR chatroom in the clear - without masking his IP address as was his normal practice.
  • Paula Broadwell, who had an affair with then CIA Director David Petraeus, was identified despite only logging into the anonymous email account created and used for the affair from public internet sites.  The FBI reportedly identified her by matching hotel and service receipt records from the times of the emails, and finding hers was the one name in common.
Schneier's point is that Internet traffic is widely tracked, and not only by governments and counter-espionage organizations.  Google does it on everything running through one or another of their sites.  Google also tracks and records websites and content for its search engines.  Blogger.com, for example, lets me know who's visited this blog, where you're from, what OS you're running, and how you found me.  Apple tracks user behaviors on iPhones and iPads.  Facebook tracks its members and their behaviors, and backs up their content and submissions.  They've also admitted tracking their members non-Facebook activities, and using cookies to track online behaviors of non-members who visit Facebook pages.  Pretty much every commercial site builds profiles of users and customers.  And metrics firms collect and track data (anonymized, they say) on users and their Internet behaviors in terms of data traffic flows..

Now if all these were separate, private, and secure, they may be seen by many as the acceptable cost for the services and benefits provided by the Internet and various online services.  Even if they were shared, it might not be so bad, if it would take significant time and effort to try to link things together (particularly if you're looking for patterns in behavior).  If "surveillance" was too costly or inconvenient to be used regularly or for trivial purposes.
However, that's increasingly not the case, due to technology advances and the rise of Big Data.  If you haven't heard the phrase before, Big Data refers to a range of programs and techniques for trolling extremely large data sets (such as online tracking data) to tease out and identify patterns and links.  With Big Data to help, the sheer volume of online data is no hindrance.  Automated systems can scan millions of emails in real time looking for key words or phrases.  Automated systems can match online searches, or the use of certain apps, to purchasing behaviors and location data from mobile devices to send users a coupon for a nearby store or restaurant.  And data storage costs keep falling.  (And while not exclusively Internet, facial recognition software and the myriad private and public video cameras can be used to track a person's movements).
  As Schneier puts it,
This is ubiquitous surveillance: All of us being watched, all the time, and that data being stored forever. This is what a surveillance state looks like, and it's efficient beyond the wildest dreams of George Orwell.
Nor does there seem to be an easy solution, or a means of opting out.
There are simply too many ways to be tracked. The Internet, e-mail, cell phones, web browsers, social networking sites, search engines: these have become necessities, and it's fanciful to expect people to simply refuse to use them just because they don't like the spying, especially since the full extent of such spying is deliberately hidden from us and there are few alternatives being marketed by companies that don't spy.
So, Schneier concludes, welcome to an Internet without privacy; welcome to the Internet surveillance state.  While public interest groups try to raise concerns about privacy, and individuals rant, the public doesn't seem to mind - as long as Amazon and Netflix make good recommendations, YouTube lets you know about the latest "cute kitty" viral video, and social media don't charge fees.  The Internet was never truly private in the first place, and isn't likely to ever significantly shift in that direction.  In part because one of the significant public values of the Internet comes from the lack of privacy and the ability to find and make connections.  What international and national regulatory moves there are are about giving governments more control over the Internet, and more access to the information it transmits and generates.  Which means even fewer real privacy protections.

  If that worries you - and it should - you could go offline.  But in a modern global information society, that comes at a high cost.  Or you could try to level the playing field, as David Brin suggests in The Transparent Society - let us, as citizens, have the same access to surveillance of government activities as the government has over our activities.  Make government truly transparent, rather than settling for "transparency" being defined as giving people access to information the government wants to provide them.  Turn the cameras around, open records, and let the public see what government actually does, rather than only what the government claims it's doing (true or not).  Or hoping that an (increasingly scarce) honest and aggressive press will investigate and report, and do the monitoring for them.

Sources -  The Internet is a surveillance stateCNN Opinion
David Brin's Transparency website

Edit - fixed some typos

Tuesday, November 20, 2012

Leahy's Email Privacy Bill does 180

Last May, Senator Pat Leahy (D, Vt.) introduced the "Electronic Communications Privacy Act Amendments Act of 2011" (PDF), ostensibly to ensure that police and government agencies needed a search warrant to access private conversations and information about locations of mobile devices.  The bill was, in part, a response to arguments from Obama's Dept. of Justice that "warrantless tracking should be permitted because Americans enjoy no "reasonable expectation of privacy" in their, or at least their cell phones', previous locations."

As the bill nears its scheduled vote next week, its come out that the bill has been dramatically rewritten.  Rather than protecting privacy, the revised bill specifically allows more than 22 Federal agencies "to access Americans' e-mail, Google Docs files, Facebook wall posts, and Twitter direct messages without a search warrant."  The revised bill would also expand the powers of the FBI and Homeland Security "to gain full access to Internet accounts without notifying either the owner or a judge."
Christopher Calabrese, legislative counsel for the American Civil Liberties Union, said requiring warrantless access to Americans' data "undercuts" the purpose of Leahy's original proposal. "We believe a warrant is the appropriate standard for any contents," he said. 
Leahy is said to have been pressured by groups representing police and district attorneys, as well as some strong politicking from the US. Justice Dept., to limit online privacy protections, or at least include major exemptions.  The revised bill does retain some protections from local and state police actions, but critics say it opens the floodgates for misuse at the Federal level.

Source -  Senate bill rewrite lets feds read your e-mail without warrantsCNet News

Saturday, September 29, 2012

Interpreting COPPA and Children's Privacy - A Step Too Far?

The main thrust of the Children's Online Privacy Protection Act (COPPA) was to limit Web site operators collecting personal information from children without the express permission of their parents.   The FTC (Federal Trade Commission), which oversees COPPA compliance, recently started considering expanding COPPA coverage and reach by broadening definitions of "personal information," "knowingly collecting," and website "operator."  While COPPA explicitly gives the FTC the ability to define, or re-define, those terms, the proposed expansions are substantial, and would significantly expand both the activities covered and the online sites and services covered.  While a number of children's and privacy advocacy groups fully support efforts to protect children's privacy, a number of industry groups have legitimate concerns that some of the proposed expansions could have significant unintended effects for all Internet users.  The Interactive Advertising Bureau (IAB), for example, indicated in formal comments to the FTC that enforcement of the expanded definitions (as formally proposed) could "restrict children’s access to online resources by undermining the prevailing business model" and "pose technical challenges to the effective functioning of the online ecosystem."
  So what are the issues, and what's really at stake?

"Operators" - or who does COPPA apply to? 
The statutory language of COPPA defined operators essentially as commercial website operators who collect personal information from users and where the website is either directed towards kids, or are general interest sites that knowingly collect personal information from children under the age of 13.  The current proposals would ad to that group third-party services such as social media plug-ins, ad networks, online gaming, and mobile apps.  In proposing the expansion, the FTC provided this rationale -
"The Commission now believes that the most effective way to implement the intent of Congress is to hold both the child-directed site or service and the information-collecting site or service responsible as covered co-operators... (A)n operator of a child-directed site or service that chooses to integrate into its site or service other services that collect personal information from its visitors should be considered a covered operator... Although the child-directed site or service does not own, control, or have access to the information collected, the personal information is collected on its behalf."
While the intent may be to assure that sites don't avoid protecting kids' privacy by farming data collection to a third party, it's difficult to frame regulatory language that would differentiate websites take use third parties to collect personal information, and/or benefit from collected user information, from websites with no interest in, or use for, user data yet link to services and sites that do. The proposed expansion would also make third party operators who collect user information liable for COPPA compliance if any affiliated or networked website is directed towards children, regardless of whether the "third party operator" has any interest in, or intent to, collect user information from children. If the implementing language is too broad, it could have the effect of making every website or online service provider legally liable for the content focus and user data collection practices of every website or online service they are linked to, or interact with. Given the nature of the internet, holding publishers liable for the COPPA compliance of affiliated services or linked sites and services would likely create a logistical nightmare of previewing and vetting of content, focus, and any user data collection practices. In their formal comments on the proposed changes, the Interactive Advertising Board (IAB) claimed that "pose technical challenges to the effective functioning of the online ecosystem." Particularly for website operators or online publishers who aren't commercial and have no interest in, or use for, user data.

"Knowingly collecting," or what evidence of intent or purpose is required?
The FTC is also proposing to expand the standard of intent by shifting from applying to operators who knowingly collect kids' personal information, to apply when operators might have "reason to know" that personal information is being collected from or content is directed towards children under the age of 13.
In regulatory and legal circles, reason-to-know is widely acknowledged as a broader, looser standard than actual knowledge of actions or behaviors. The FTC, as noted in the quote above, sees their mandated purpose as protecting children's privacy by requiring parental consent to collect personal information from kids, even if it is not knowingly and intentionally collected. The reason-to-know standard would extend COPPA to at least some incidental collection of covered user data, but not the absolute coverage that many advocacy groups have called for. They would prefer to see that privacy coverage and requirements for parental consent for data collection from kids be universal, to assure that no personal information is ever collected from children under thirteen without explicit parental consent.
Implementing a vaguer and looser standard can be problematic - "knowingly" is a clear and precise standard, even if can be difficult to provide. "Reason-to-know" is not precise, but has been interpreted in other settings as existing when an individual could reasonably expect something is probable - in this setting, would not be surprised if a third party operator collected personal user information or directed content or services to kids. Still, there's a lot of imprecision and uncertainty left - for example, should a blogger targeting seniors that links to an online social gaming app be e know what user information the app collects, or whether children under 13 are playing that social game app?
And if combined with an expansion if the definition of operators to second and third parties, the costs of compliance are spread to those who are only peripherally involved with children or collection of user data.

"Personal information," or just how personal does information need to be?
To a very large extend, the Internet, mobile, and social media systems run on user data, because sending and receiving information requires some kind of address. Data transfers online need IP addresses; mobile communications require unique identifiers for devices or users; and social media need to know where to send whatever stuff we share with friends and followers. The original statutory language of COPPA used older offline definitions widely used in privacy contexts - names, street addresses, social security numbers, phone numbers; and added email addresses as a nod to the Online context. But in an ever-evolving online ecosystem, these aren't our only addresses, or unique identifiers. If the concern about collecting personal information is that whatever other information or behaviors that are being collected can be directly linked to a specific individual, then the FTC really does need to look at what it defines as personal information.
Last year, the FTC proposed expanding the definition of "personal information" to include any "unique identifier" that could be used to link a child's activities on multiple sites. The proposal identified a few examples of unique identifiers - IP addresses, device serial numbers, tracking cookies. The online world is replete with unique identifiers; as are the worlds of mobile devices, wireless services, mobile phones, and social media. As I said earlier, they all need addresses - and addresses that aren't relatively unique identifiers aren't that useful. Are the FTC's examples appropriate?
In one sense, clearly not. As the IAB pointed out, the problem with the listed identifiers is that they aren't necessarily user-specific - what they are are primarily device identifiers. If there are, or may be, multiple users, that can decouple these unique identifiers from an unique person. (We've gone through this with IP addresses, which were initially permanently assigned to a device. When the number of devices exploded, and Internet Service Providers noted they weren't always on, they switched to dynamic IP addressing, where the unique address is assigned when the device is actively connected, but tossed back into the ISP's pool of IP addresses when the device was disconnected, to be assigned to another device when it actively connects. To uniquely link an IP address with a specific computer, you now need both the dynamic IP address and the time). The proposed new unique identifiers permit the delivery of content and advertising to a device, not to an identified individual," the IAB argues.
In addition, device identifiers are largely automatically generated and provided with online activities without user input or direct authorization. This creates a variety of potential issues - are dynamic IP addresses new unique identifiers that require user or parental validation of permission to use? would COPPA be invoked if several distinct online services share a common password/login (linking across sites)? Would Internet-connected devices need to be child-proofed in the absence of parental consent to collecting device identifiers? How might this impact "TV Everywhere" implementation, which needs unique identifiers not only as device address, but for validation of eligibility to receive specific content? How might that affect the potential distribution of children's programming, or educational content or games? There's a real conflict between the need for tracking use and validating eligibility through the use of unique identifiers and tracking user behaviors and the primary funding mechanisms for websites and online services (advertising and subscriptions). Defining unique identifiers poorly or inappropriately would create significant compliance costs that could only be avoided by prohibiting children's access and use. In such a case, the IAB expressed concern that it could "restrict children’s access to online resources by undermining the prevailing business model."
A closer look at the FTC's proposals and supporting arguments suggests that their real concern was the potential use of behavioral advertising techniques on children under 13. The FTC did include a specific proposal for a ban on using behavioral targeting techniques on young children without their parents' permission. But the courts can be reluctant to apply content-related bans without specific evidence of harm. That could explain the FTC's choice of specific unique identifiers and emphasis on linking behaviors and information across sites - their list mirrors what is needed for behavioral advertising to occur. Thus, the FTC may have felt that expanding the definition of "personal information" in that specific direction could be a backdoor means to limit behavioral advertising to kids. The problem here is that these same elements are also at the heart of a great many other online services and activities, so this expansion would have unintended (I hope) negative consequences in many other areas. Particularly if the expansion of "personal information" to include a range of other "unique identifiers" and the idea of "persistent identifiers" defined as identifiers shared across sites or services, gets carried through to other privacy regulation.
Including device registration numbers as "personal information" could really impact the rapidly expanding growth of mobile services, as apps and services would need to find other means to identify and validate devices and uses. The whole foundation of social media and interconnected sites and services is similarly built on the availability of "persistent identifiers."

A Step Too Far?
The FTC clearly has the authority to consider redefining these key aspects of COPPA, and strong arguments can be made that it needs to, considering how the online world has changed in the last decade. (Not to mention the pressures being applied by a variety of advocacy and industry groups).  The most immediate need is for the FTC to seriously consider expanding the definition of "personal identifiers."  The original statutory examples are mostly borrowed from regulatory language applying to analogue and physical concerns.  The language, for the most part, is far too narrow to reflect data or information that can identify individuals in an online world filled with myriad "unique identifiers" that could easily be used to link individuals with the information they provide and the actions they take online. But you can't ban or limit the use of all unique identifiers without crippling the Internet, or an increasing number of media devices and services - or banning their use by the people who's privacy you're trying to protect. Redefining "personal information" needs to be approached with a surgeon's scalpel rather than a blunderbuss, as any change is likely to have widespread and profound implications.
  In any consideration of expanding the kinds of identifiers to be included in a definition of "personal information" the FTC (and regulators generally) shouldn't pick them because they might achieve a specific policy goal. Even if they do, they'll also impact any other uses that rely on or utilize that specific type of identifier. Regulators need to consider the other implications and effects of proposed regulatory changes before redefining things - otherwise someone's likely to wonder why it didn't do what it was supposed to, and/or how to fix the mess it's created somewhere else.

Sources  -  FTC Proposes New Curbs On Collecting Data From ChildrenOnlineMediaDaily
IAB: Proposed Children's Privacy Rules Undermine Business Model,  OnlineMediaDaily
FTC,  Proposed Rules Changes for Children's Online Privacy Protection Rule
FTC's COPPA website










Friday, July 13, 2012

Tracking Social Media chatter over News Events

  One of the fun things for researchers is the huge amount of public data available online, and particularly the ability to capture snippets of public opinion and reactions through social media.  (Although the methodologist in me needs to remind everyone that social media content is not necessarily representative of general public opinion.)  For example, here's some summaries of reaction to the Supreme Court's decision on the Affordable Care Act (Obamacare).

  In addition to tracking social media, many search engines provide the capacity to track trends in the use of search terms.  Google, for instance, provides access to current trends as well as the ability to track term usage over time (or at least back to 2004) for general usage, as well as in terms of news references.  Yahoo, Bing, AOL, Twitter, and YouTube all offer similar services.  Many of these also can provide user information in aggregate, and can be further broken down by geographic location.

Online metrics firm Crimson Hexagon recently provided a demonstration of some of its Online traffic tracking services though an analysis of the reaction to the Supreme Court's health care ruling.  They posted some screenshots of their real-time analytic platform ForSight that provided both descriptive summaries and more qualitative analysis of themes expressed through social media.  






   As I said at the beginning, the online world's providing a lot of interesting ways to quickly capture and analyse public opinion, comments, and reactions.  Get to work, folks.

Source - Healthcare Ruling Sparks 12,000 Tweets Per Minute, Mashable
What People Search For - Most Popular Keywords, Search Engine Watch
Obamacare: Online Reaction to SCOTUS Ruling (Update)  Crimson Hexagon press release