In the late 2000s, just a few years after the iPhone hit stores, ICSI researchers noticed something alarming in a public dataset they were using for a machine learning project: based on geotags associated with most photos and videos taken with a smartphone, it would be shockingly easy to infer a person’s physical location in the real world. A clip of a toddler playing in her backyard could be used to pinpoint her address, for example, while vacation photos could let thieves know when a home is likely unoccupied. Most smartphone users had no idea such tags were included in the content they were sharing online. 

Gerald Friedland (then at ICSI, now at Amazon Web Services) and colleagues coined a term for the potential misuse of geotags in data—“cybercasing”—and shared their concerns about this privacy vulnerability at HotSec’10 and NSPW’11. Media outlets quickly picked up the story, helping to raise the visibility of the issue and alert smartphone users to the risks of posting geotagged content on sites like YouTube, Twitter, and Craigslist. One of the team’s proposed solutions—to use geotags that are less precise, centering on a general area rather than the exact location of the phone—is now standard for many apps and operating systems. This helps reduce risks to some extent, but it remains essential for users to be aware of just how much information they are sharing when posting a photo or video. To that end, Friedland, Serge Egelman, and other collaborators from ICSI and UC Berkeley created a curriculum for teaching high school students about online privacy and launched an app, “Ready or Not?” which gave users a first-hand look at what social media posts can reveal about a person’s real-world activity. 

In parallel with these projects, Friedland and colleagues continued the machine learning work that had brought the privacy issue to their attention in the first place. After several years of negotiation, the researchers secured permission from Flickr to extract 100 million photos and videos from the site’s Creative-Commons licensed content. The collection, called YFCC100M, proved a powerful resource for machine learning research and experimentation in geo-analysis, computer vision, and other multimedia studies. Even as it was valued for its contributions to the field of computer science, this resource and its use by academic researchers and companies also raised concerns about the ethics of AI-based facial recognition systems and drew attention to the challenges of “unsharing” content once it’s been made public—further underscoring the complexities of online privacy. 

What made ICSI a good place to pursue these projects?

“These are projects that would only happen at ICSI. At ICSI, we were able to go after what just made sense. At a company, you’re often on this product side or that product side, not both. And even if you’re faculty at a university, you’re often a machine learning person or a privacy person, but not both. Being able to investigate all sides of this allowed us to get a view that nobody else has. And we were way ahead of the curve.” 


Gerald Friedland

Former Postdoctoral Research Scientist; now Principal Scientist at Amazon Web Services

This story was published in January 2026 as part of a retrospective series highlighting ICSI’s accomplishments and impacts over the years. To learn about our ongoing work, explore our Core Research Themes.