Mechanize builds reinforcement learning environments that frontier AI labs use to train and evaluate their coding models. Learn more at mechanize.work.
AI models have gotten good at narrow coding tasks but still fail at the complex, judgment-heavy parts of software engineering. We build the environments that expose those failures and help models improve.
You'll design, build, and refine RL tasks. Each task is a self-contained software engineering challenge with a prompt, an environment, and an automated grader. You own the full lifecycle: coming up with the idea, implementing the grading infrastructure, running frontier models against the task, analyzing where and why they fail, and iterating until the task is rigorous and fair.
Coming up with good task ideas requires being clever: finding situations where a frontier model will fail in interesting ways, which means seeing gaps that the model itself doesn't see. You will use coding agents heavily, and a large part of the job is directing them well, evaluating their output, and knowing when they are failing in subtle ways.
Strong technical fundamentals combined with an intuition for AI model behavior. You need to anticipate where a model will take shortcuts, distinguish genuine capability gaps from grader issues, and understand how a model will interpret a prompt. Most engineers significantly underestimate what frontier coding agents can already do; candidates who have spent significant time working with them will have a real head start.
We're happy to hire candidates who are very capable, even if they have no prior professional software experience.
This is independent, high-ownership work. You own your tasks from start to finish, with regular check-ins and feedback.
Compensation includes a $300,000 base salary, equity, and performance bonuses. Top performers can earn more in bonuses than in base salary.
Strong performers are recognized and promoted quickly. Benefits include health, dental, vision, and life insurance.
~20 person team in San Francisco. Backed by Patrick Collison, Nat Friedman, Daniel Gross, Jeff Dean, Dwarkesh Patel, and Sholto Douglas. Featured in the New York Times, the Dwarkesh Podcast and Hard Fork.
Learn more about the interview process:
Learn more about the work:
#J-18808-Ljbffr...Position Title: Research Analyst I (Mid-Level) Location: Remote (Must be a U.S. Citizen) Job Type: Contract/Consultancy Deadline... ...& Skills Bachelors degree in a field related to social science research and a minimum of 35 years of related experience or...
Query Talent is recruiting for an available CDL-A driver opportunity in Chicago, IL. This is a regional opportunity. Trailer type: Van... ...- Experience requirement: 6+ months CDL-A experience.- Driver type: Company Driver / Team Driver.- Trailer type: Van / Intermodal.
Job Description Job Description At Tesco Controls (a UFT company), our culture is grounded in the idea that the how we achieve results is just as important as the results themselves. Our winning behaviors guide the way we work, collaborate, and growboth as individual...
...Cole Law Group, P.C. is seeking an exceptional Legal Research Specialist to support the Firm's attorneys in complex litigation and other sophisticated legal matters. This is a research- and writing-intensive position. The successful candidate will have outstanding...
...We have an opening for a talented Aldi Stocker to perform daily responsibilities with dedication. Provide excellent interactions with customers and colleagues. Provide excellent interactions with customers and colleagues. Perks include competitive pay, flexible schedules...