testdriver-agent
How the TestDriver agent behaves on GitHub issues, pull requests, and @mentions
Detect all UI elements on screen using OmniParser
$ npx -y skills add testdriverai/testdriverai --skill testdriver-parse --agent claude-codeHow it fires
How this skill gets triggered: by you, by Claude, or both.
/testdriver-parseContext preview
The summary Claude sees to decide when to auto-load this skill.
Detect all UI elements on screen using OmniParser
name: testdriver:parse description: Detect all UI elements on screen using OmniParser
<!-- Generated from parse.mdx. DO NOT EDIT. -->
Parse the current screen using OmniParser v2 to detect all visible UI elements. Returns structured data including element types, text content, interactivity levels, and bounding box coordinates.
This method analyzes the entire screen and returns every detected element. It's useful for:
<Note> **Availability**: `parse()` requires an enterprise or self-hosted plan. It uses OmniParser v2 server-side for element detection. </Note>
const result = await testdriver.parse()
None.
`Promise<ParseResult>` - Object containing detected UI elements
| Property | Type | Description | |----------|------|-------------| | `elements` | `ParsedElement[]` | Array of detected UI elements | | `annotatedImageUrl` | `string` | URL of the annotated screenshot with bounding boxes | | `imageWidth` | `number` | Width of the analyzed screenshot | | `imageHeight` | `number` | Height of the analyzed screenshot |
| Property | Type | Description | |----------|------|-------------| | `index` | `number` | Element index | | `type` | `string` | Element type (e.g. `"text"`, `"icon"`, `"button"`) | | `content` | `string` | Text content or description of the element | | `interactivity` | `string` | Interactivity level (e.g. `"clickable"`, `"non-interactive"`) | | `bbox` | `object` | Bounding box in pixel coordinates `{x0, y0, x1, y1}` | | `boundingBox` | `object` | Bounding box as `{left, top, width, height}` |
const result = await testdriver.parse();
console.log(`Found ${result.elements.length} elements`);
result.elements.forEach((el, i) => {
console.log(`${i + 1}. [${el.type}] "${el.content}" (${el.interactivity})`);
});const result = await testdriver.parse();
const clickable = result.elements.filter(e => e.interactivity === 'clickable');
console.log(`Found ${clickable.length} clickable elements`);
clickable.forEach(el => {
console.log(`- "${el.content}" at (${el.bbox.x0}, ${el.bbox.y0})`);
});const result = await testdriver.parse();
// Find a "Submit" button
const submitBtn = result.elements.find(e =>
e.content.toLowerCase().includes('submit') && e.interactivity === 'clickable'
);
if (submitBtn) {
// Calculate center of the bounding box
const x = Math.round((submitBtn.bbox.x0 + submitBtn.bbox.x1) / 2);
const y = Math.round((submitBtn.bbox.y0 + submitBtn.bbox.y1) / 2);
await testdriver.click({ x, y });
}const result = await testdriver.parse();
// Get all text elements
const textElements = result.elements.filter(e => e.type === 'text');
textElements.forEach(e => console.log(`Text: "${e.content}"`));
// Get all icons
const icons = result.elements.filter(e => e.type === 'icon');
console.log(`Found ${icons.length} icons`);
// Get all buttons
const buttons = result.elements.filter(e => e.type === 'button');
console.log(`Found ${buttons.length} buttons`);import { describe, expect, it } from "vitest";
import { TestDriver } from "testdriverai/vitest/hooks";
describe("Login Page", () => {
it("should have expected form elements", async (context) => {
const testdriver = TestDriver(context);
await testdriver.provision.chrome({
url: 'https://myapp.com/login',
});
const result = await testdriver.parse();
// Assert expected elements exist
const textContent = result.elements.map(e => e.content.toLowerCase());
expect(textContent).toContain('email');
expect(textContent).toContain('password');
// Assert there are clickable elements
const clickable = result.elements.filter(e => e.interactivity === 'clickable');
expect(clickable.length).toBeGreaterThan(0);
});
});const result = await testdriver.parse();
result.elements.forEach(el => {
// Pixel coordinates
console.log(`Element "${el.content}":`);
console.log(` bbox: (${el.bbox.x0}, ${el.bbox.y0}) to (${el.bbox.x1}, ${el.bbox.y1})`);
console.log(` size: ${el.boundingBox.width}x${el.boundingBox.height}`);
console.log(` position: left=${el.boundingBox.left}, top=${el.boundingBox.top}`);
});const result = await testdriver.parse();
// The annotated image shows all detected elements with bounding boxes
console.log('Annotated screenshot:', result.annotatedImageUrl);
console.log(`Image dimensions: ${result.imageWidth}x${result.imageHeight}`);1. TestDriver captures a screenshot of the current screen 2. The image is sent to the TestDriver API 3. OmniParser v2 analyzes the image to detect all UI elements 4. Each element is classified by type (text, icon, button, etc.) and interactivity 5. Bounding box coordinates are returned in pixel coordinates matching the screen resolution
<Note> OmniParser detects elements visually — it works with any UI framework, native apps, and even non-standard interfaces. It does not rely on DOM or accessibility trees. </Note>
<AccordionGroup> <Accordion title="Use find() for targeting specific elements"> For locating and interacting with a specific element, prefer `find()` which uses AI vision. Use `parse()` when you need a complete inventory of all elements on screen.
// Prefer this for clicking a specific element
await testdriver.find("Submit button").click();Repo: testdriverai/testdriverai
How the TestDriver agent behaves on GitHub issues, pull requests, and @mentions
Deploy TestDriver on your AWS infrastructure using CloudFormation
How TestDriver learns your app and caches what it discovers for instant, deterministic replays