شرح موقعیت
Abu Dhabi (UAE) · Member of Technical Staff · Engineer II, L4 · MJD02.0.1 Hands-on engineer in the R&D team building AI agents that control web browsers with open-weight LLMs. You sit with the domain experts who operate legacy web applications, learn how they actually work, and turn it into what the agents are trained and tested against. Greenfield. THE TECHNICAL CHALLENGE We build agents on open-weight LLMs that operate web applications through the browser. The applications were built for people, and the people who operate them work largely by intuition; a documented process is the exception. This role is deployed into that reality on behalf of the R&D team, with no finished product to install: sit with domain experts, record how they work, and turn what you see into structure that has to hold. Workflow models. Evaluation tasks from a single action to a whole task. Fault specifications for simulated copies of the applications. Demonstration recordings. What fails on real applications comes back to the team as a specification. Progress is measured against benchmarked results. KEY RESPONSIBILITIES Observe and record how domain experts operate their applications, and elicit what the recordings do not show Decompose workflows into explicit structure: states, decisions, exceptions and failure modes, generalised across applications Turn that structure into artefacts that must work: evaluation tasks, fault specifications for simulated applications, demonstration recordings Run the feedback loop from real applications to the R&D team; build quick prototypes and demonstrations DESIRED QUALIFICATIONS Has worked embedded with expert users on behalf of an engineering or research team, and can say what the users could not tell them Has turned a messy, undocumented human process into a specification that a team implemented and that held in production Has built evaluation tasks or simulated environments from real workflows Has recorded and analysed how third-party web applications behave (network logs, DevTools, screen recordings) to diagnose failures EXPECTED QUALIFICATIONS Patient and precise with people who have no documented process; runs discovery sessions with non-technical experts Thinks in states, decisions and exceptions, and treats the happy path as the easy part Builds a working prototype in days in TypeScript or Python; has used browser automation (Playwright or comparable) against real websites Daily use of AI coding tools, including reviewing and verifying their output Presents working demonstrations to non-technical audiences —— HOW WE WORK Product engineering: we own what we build and run it in production Small teams, two-week cycles, working software at every review AI coding tools are part of the standard workflow PROCESS In coding exercises, AI tools are allowed and expected. No LeetCode. WHO WE ARE New product organisation as part of a large semi-government in Abu Dhabi. International, ex-FAANG team. Completely greenfield, with a modern tech stack. —— REQUIREMENTS TO BE CONSIDERED RELATED TECHNOLOGIES AND CONCEPTS
مزایا
- Founding-team scope
- AI-augmented engineering environment
- Access to on-premise Nvidia B200s
- Flexible work environment
- Introductory call
- Technical conversation
- Practical session; the format is agreed with you
- Fluent Arabic and English, spoken and written
- 3+ years in an engineering role with direct contact with the people who use what you build
- Own code operated in production; comfortable in TypeScript or Python
- Bachelor's degree in any field, or self-taught with a track record of open-source contributions
- Browser automation and computer use: Playwright, Chrome DevTools Protocol, Stagehand, Browser Use, Steel Browser
- Session capture and replay: rrweb, HAR, WARC, screen recording
- Evaluation: evaluation harnesses, trajectory evaluation, LLM-as-judge, Inspect, Langfuse; benchmarks such as Online-Mind2Web, WebArena, OSWorld
- Process analysis: task mining, process mining, cognitive task analysis, behaviour-driven specification (Gherkin), state machines
- Agent frameworks: OpenHands, SWE-agent, Browser Use, LangGraph, smolagents
- Models and infrastructure: open-weight LLMs, vLLM, SGLang, Kubernetes, TypeScript, Python