# An example of why Parallel beat Exa in our AX eval

A shortened preview was only the start. Parallel gave the agent a concrete way to inspect the saved response and keep researching.

Published: 2026-10-09
By [Duc Ngo](https://www.linkedin.com/in/dwcngo)

<p className="writing-lead"><strong>Parallel handled the long-document task better in our AX eval.</strong> Its Skill explained how to continue after a shortened preview: <strong>save the response, open the file and search it for the relevant passages.</strong> The agent followed it and passed.</p>

Explore the [**full comparison**](https://reports.malmhq.com/evaluations/comparisons/search-research-suite-2026-10-04/results.html) and public interaction logs.

## The job: check a claim about DeepSeek

We asked an agent to investigate a viral claim: DeepSeek R1 matched OpenAI o1 at a fraction of the API price. It had to produce a sourced memo and evidence tables covering benchmarks, prices and training costs.

Both attempts used Codex with GPT-6 Luna. Both found the R1 paper. Their next actions differed.

<GuidanceDiagram />

## Parallel gave the agent a way to keep reading

A Skill provides instructions for using an interface. Parallel’s explained how to **inspect a saved JSON response** when the **terminal preview was shortened**.

The agent **saved the returned paper text in a local JSON file**. It used Python to **search for Table 4 and key headings**, then print the surrounding passages. This let it read the benchmark details and later limitations without putting the entire file into the conversation. The response still had a 50,000-character limit; the file preserved text beyond the terminal preview. Its memo said the retrieved passages did not provide a comparable total R1 training budget. It **stayed within the inspected evidence**. The trial passed.

## Exa’s agent stopped too early

Exa also supplied instructions. They said to read the returned text directly in the conversation and avoid local file-processing tools. They also said to **avoid character limits**. The agent instead **capped its PDF response at 18,000 characters**. It stopped during section 2. The agent wrote its memo **without retrieving later sections**.

Its memo then said the paper did not give a complete R1 training-cost total. **Partial text could not support that whole-paper claim.** The eval system recorded this as failed. The agent ignored guidance; Exa had tools to continue.

Both judged the viral claim misleading. Their source support differed.

## Why good Skill design matters for AX

**A good Skill gives the agent a practical way to finish the job.** It explains which action to take, what to expect from the response and how to continue when information is missing or cut short. Concrete steps and examples make the instructions easier to use.

Both providers supplied Skills. **Parallel’s guidance was better for this situation:** it explained how to save the returned content and inspect the saved file when the preview was shortened. The agent then searched for relevant headings and read the matching passages. Exa advised reading inline and avoiding character limits, but its agent capped the response and stopped early.

For large documents, **design the Skill around the reading workflow**: where the content goes, how to find the relevant parts and how to check that enough evidence has been inspected before writing a conclusion. Apply the same approach to missing information and failed requests: give the agent a specific next action it can actually carry out.

## Give the agent a usable next step

Parallel’s guidance was better here. It made **the next action concrete**.

Customers bring different agent runtimes. You define available actions, their requirements, human approval boundaries and call records, including refusals. Give agents practical instructions for continuing when a response is shortened or a request fails.

We ran these evals with Intermesh. See the [Parallel report](https://reports.malmhq.com/evaluations/parallel-2026-10-04T151004081Z/report.html) and [Exa report](https://reports.malmhq.com/evaluations/exa-2026-10-04T151004081Z/report.html) and the [**full comparison**](https://reports.malmhq.com/evaluations/comparisons/search-research-suite-2026-10-04/results.html).
