AI
Perplexity Open Sources WANDR Benchmark for AI Research Agents
Perplexity AI released WANDR, an internal benchmark for evaluating deep and wide research capabilities, as open source.
Perplexity AI, the artificial intelligence powered search engine company, announced on Monday, July 27, 2026, that it has open sourced WANDR, an internal benchmark it built and used to develop deep and wide research capabilities within its flagship product, Perplexity Computer. The announcement was made via Perplexity AI's official account on X (formerly Twitter). In the post, the company stated, "We're open sourcing WANDR. WANDR is an internal benchmark we built and used for building deep and wide research capabilities inside Perplexity Computer." The post also included a link to a technical paper detailing the benchmark. According to the technical paper published at research.perplexity.ai, WANDR stands for Wide and Deep Research. The benchmark is designed to evaluate research agents that must search both broadly and deeply.
The paper explains that wide and deep research requires two primary capabilities. First, the agent must search broadly enough to find all qualifying entities. Second, the agent must investigate deeply enough to support each claim with evidence. WANDR represents these requirements as hierarchically verifiable records. The benchmark consists of 500 research tasks, each designed to test an agent's ability to balance broad exploration with deep investigation. Each task produces a record that can be independently verified, enabling more granular evaluation of agent performance. Perplexity's decision to open source WANDR marks a significant step in the open AI ecosystem. By releasing this benchmark, Perplexity not only provides an evaluation tool for the research community but also allows other developers to test and compare their own research agents.
This has the potential to accelerate innovation in developing AI agents capable of complex research. Perplexity Computer, the product mentioned in the announcement, is a platform that enables users to conduct deep research with AI assistance. WANDR was used internally to ensure that agents within Perplexity Computer can handle tasks requiring information retrieval from multiple sources and cross verification. By open sourcing this benchmark, Perplexity invites contributions from the community to improve and expand WANDR's capabilities. This move can also be seen as an effort by Perplexity to establish an industry standard for evaluating AI research agents. Currently, there is no widely accepted benchmark for measuring an agent's ability to conduct wide and deep research. WANDR fills that gap by providing a systematic and reproducible framework.
The WANDR technical paper explains that the benchmark is designed to address the limitations of existing benchmarks, which tend to focus on a single dimension, such as depth or breadth, but not both. WANDR combines both dimensions into a single evaluation framework, providing a more holistic view of agent capabilities. Each task in WANDR is designed to simulate real world research scenarios, where an agent must navigate large amounts of information, identify relevant sources, and construct evidence supported arguments. Evaluation results can then be used to identify specific weaknesses in an agent and guide further development. Perplexity has not announced a detailed timeline for releasing the WANDR source code, but the open source announcement indicates that the code will be available in the near future.
The AI developer and researcher community is expected to welcome this move, as it provides a much needed tool for comparing and improving research agents. By releasing WANDR as open source, Perplexity also strengthens its position as a player that supports transparency and collaboration in AI development. This step aligns with trends in the AI industry, where companies like Meta and Google have also released some of their models and tools as open source. Looking ahead, WANDR has the potential to become a standard benchmark for evaluating AI research agents, similar to the role of GLUE and SuperGLUE in natural language understanding.
If widely adopted, this benchmark could drive the development of more sophisticated and reliable agents, which in turn would improve the quality of AI assisted research. Perplexity's open source release of WANDR is a strategic move that not only benefits the broader AI community but also positions the company as a leader in transparent AI development. The benchmark's focus on both breadth and depth addresses a critical gap in current evaluation methods, and its hierarchical verification system ensures rigorous assessment. As AI research agents become more prevalent, having a standardized evaluation tool like WANDR will be essential for measuring progress and fostering competition. The 500 task dataset provides a robust foundation for testing, and the open source nature allows for continuous improvement through community contributions.
This initiative could spur further research into multi dimensional agent evaluation, ultimately leading to more capable and trustworthy AI systems.