A short video shows an ordinary hotel room. Nothing identifies the building, city, or even the country. Its location metadata has been stripped away. But the quiet sound of traffic outside, the shape of the electrical outlets, the room’s acoustics, and a few objects visible in the background may still reveal exactly where it was recorded.
The U.S. Intelligence Community wants artificial intelligence (AI) that can assemble those fragments and locate the camera within roughly 820 feet in less than five minutes.
The Intelligence Advanced Research Projects Activity, better known as IARPA, quietly released the program nearly five months ago. Called LocUS, or Location Using Sound, the program seeks automated systems that can determine where previously unseen videos, images, or audio clips were captured without relying on embedded metadata.
If successful, the technology could help intelligence and law enforcement analysts locate hostages, identify sites connected to human trafficking, or trace videos posted online by hostile actors. But its capacity to extract a location from seemingly incidental sights and sounds would also demonstrate how difficult truly anonymous multimedia has become.
“These research programs will help build capabilities that are directly applicable to mission needs by bridging the technical gap between emerging solutions and successful application,” Principal Deputy Director of National Intelligence Aaron Lukas said when IARPA released LocUS and four other programs at an April Proposers’ Day meeting. “By working together, we can achieve great things and ensure our nation remains secure and prosperous in the face of emerging threats.”
IARPA describes LocUS as an intelligence-grade extension of the same general problem GeoGuessr players face. This popular game asks users to identify locations from Google Street View images.
Experienced players study vegetation, road markings, utility poles, architecture, and other small details. Researchers have since applied large language models, computer vision, and automated search tools to outperform some of the game’s best human competitors.
LocUS raises the difficulty considerably. The system must work indoors and outdoors, across search areas ranging from individual neighborhoods to entire continents. It must also process degraded material, including blurry video, low-quality images, and audio with limited bit rate or sample quality.
Most importantly, it cannot simply be a better image-matching engine. IARPA requires systems to combine visual and audio information, including environmental sounds that may carry geographic clues even when no useful speech is present.
According to the agency’s draft technical summary, the initial target is for 90 percent of predictions to fall within 820 feet, or 250 meters, of the true location. A system searching an area as large as a continent, defined in the program materials as approximately 1,550 miles across, would have less than five minutes to return an answer.
The software must operate autonomously. Human analysts cannot guide the system through a query, and IARPA will test submitted technologies using previously unseen multimedia held in sequestered datasets.
That testing requirement is significant because geolocation models can appear remarkably capable when they recognize places represented in their training data. Performance becomes much harder to sustain in unfamiliar locations, particularly indoors, where hotel rooms, offices, and apartment buildings may share nearly identical designs.
Public responses to IARPA’s Proposers’ Day teaming request show that researchers were already organizing around that problem. The list does not identify selected contractors, but it offers an unusual look at the technologies and datasets that prospective participants believed LocUS would require.
Dr. Abby Stylianou, a computer science professor at Saint Louis University, featured her laboratory’s work with the National Center for Missing and Exploited Children. The team helped develop TraffickCam, an application that allows travelers to upload photographs of hotel rooms for use in trafficking investigations.
In a response to IARPA, Dr. Stylianou said submissions to TraffickCam, combined with images collected from the web, have created a massive dataset containing millions of location-specific photographs.
“It [TraffickCam] has been downloaded by over 250,000 users and receives between typically on the order of hundreds of new images each day,” Dr. Stylianou writes. “Combined with web-scraped imagery, we have a dataset of over 15m hotel and rental property images from around the world.”
Other prospective partners included researchers who previously worked on IARPA’s Finder program, which ran from 2011 to 2016 and focused on geolocating outdoor imagery. Several emphasized acoustic scene analysis, multimodal reasoning, and methods for locating images taken in places absent from a model’s training data.
Dr. Xiaoming Liu, a computer vision researcher at the University of North Carolina at Chapel Hill, identified scarcity as a major challenge. A model may have examples from every state, he noted, while still lacking useful training material from thousands of counties, towns, and individual buildings.
“You may have training data from every single US state, but what about every county within a state, or every town within a county, etc.,” Dr. Liu asks. “How well we could address this will also have a large impact to the generalization of our solution.”
LocUS is explicitly designed to stress-test those weaknesses. The planned 15-month program calls for five independent test events, approximately one every three months. Before each evaluation, research teams must provide containerized software and source code that government evaluators can install, run, and retrain.
The agency will separately measure image-only and audio-only performance to determine what each modality contributes. However, a complete LocUS system must fuse all available information into a single location prediction.
IARPA has ruled out facial recognition, voice recognition, and other biometric identification methods. Systems limited to human speech, one geographic region, or a single landscape type are also excluded.
The agency strongly prefers, but does not require, technology that can explain its conclusions in plain language. For example, identifying a particular architectural feature or background sound that supported or eliminated a candidate location.
Ultimately, that could make LocUS more useful than an automated map pin alone. An analyst could review the evidence behind a prediction, judge whether the system noticed something meaningful, and decide how much confidence to place in its answer.
The competition for LocUS performers closed in June 2026, and IARPA planned an aggressive 15-month development phase involving five rounds of testing. Although IARPA has not publicly disclosed a kickoff date or announced the selected teams, development is very likely already underway. If work began shortly after the solicitation closed, the research phase would likely continue into late 2027.
How much the public ultimately learns about its progress is less certain. IARPA may disclose research findings, publications, or broad program results, but the working status of intelligence technologies and whether they are adopted for classified missions often remains undisclosed. LocUS could therefore become an active intelligence capability without its final performance or deployment ever being fully known outside the Intelligence Community.
“LocUS will improve the geolocation capabilities of the Intelligence Community (IC) considerably beyond imagery-only methods and thereby increase the volume of content that can be accurately geolocated,” IARPA writes in a brief program overview. “Applications for national security include human trafficking interdiction, hostage recovery, and other intelligence and law enforcement use cases.”
Tim McMillan is a retired law enforcement executive, investigative reporter and co-founder of The Debrief. His writing typically focuses on defense, national security, the Intelligence Community and topics related to psychology. You can follow Tim on Twitter: @LtTimMcMillan. Tim can be reached by email: tim@thedebrief.org or through encrypted email: LtTimMcMillan@protonmail.com
