Senior Software Engineer – Python (LLM Evaluation & Repository Validation)
Senior Software Engineer – Python (LLM Evaluation & Repository Validation) is an open role at Turing AI Advancement Work, available remotely.
Applications open on Turing AI Advancement Work’s own site. How we make money.
About this role
About Turing Turing is one of the world's fastest-growing AI companies, accelerating the advancement and deployment of powerful AI systems. Founded in 2018, Turing partners with leading AI labs and enterprises to improve frontier AI models through high-quality data, expert human feedback, and rigorous evaluation frameworks. Our contributors work on cutting-edge projects spanning coding, reasoning, agentic workflows, software engineering, and multimodal AI—helping shape the next generation of intelligent systems used by millions worldwide. About the Role We are looking for experienced Software Engineers with a strong background in developing and operating large-scale production systems. The ideal candidate has worked extensively on enterprise software, understands how production systems fail, has experience identifying security vulnerabilities in code, and has successfully implemented secure fixes and mitigations. This role requires practical experience in debugging complex production issues, participating in incident response, and building resilient, secure software. What You'll Do Design and build realistic coding-agent benchmark tasks using production-like repositories, tests, configs, documentation, and runtime scenarios. Create benign engineering tasks such as bug fixes, feature additions, CI repairs, config migrations, integration updates, and runtime-state fixes. Define clear utility requirements to verify that the coding agent completes the requested task correctly. Define safety/alignment requirements to ensure the agent preserves system constraints, developer intent, data integrity, privacy, permissions, and oversight mechanisms. Develop visible and hidden test suites to evaluate both task completion and unsafe agent behavior. Create safe reference solutions and unsafe reference solutions where the agent passes utility checks but violates multiple safety/alignment requirements. Identify and evaluate unsafe shortcuts such as disabling tests, weakening assertions, deleting protected data, leaking secrets, broadening permissions, or bypassing validation. Package tasks with prompts, metadata, runnable repos, evaluators, reference patches, scoring rubrics, and calibration notes. Analyze rollouts from frontier coding agents such as Claude Code and Codex to assess utility completion and safety/alignment violations. Collaborate with engineering, QA, security, and client teams to improve task quality, evaluator reliability, and benchmark difficulty. Required Qualifications Bachelor's or Master's degree in Computer Science or a related technical discipline. Minimum 8 years of hands-on software engineering experience in leading product companies or technology startups. Strong programming experience in Python (mandatory) along with one or more programming languages such as Java, Go, C++ Demonstrated experience building and maintaining production software used by real customers. Strong understanding of software architecture, debugging, and performance optimization. Required Technical Experience Production engineering Experience developing software deployed in production at scale. Strong understanding of software deployment pipelines, monitoring, logging, and observability. Experience diagnosing production failures using logs, metrics, and distributed tracing. Knowledge of high availability, fault tolerance, scalability, and disaster recovery principles. Secure Software Development Experience identifying security vulnerabilities through code reviews or security assessments. Practical knowledge of common software vulnerabilities, including: Injection attacks Authentication and authorization flaws Memory safety issues Race conditions Deserialization vulnerabilities Secrets management issues Input validation failures Experience implementing secure code fixes and validating remediation. Perks of Freelancing with Turing: Work in a fully remote environment. Opportunity to work on cutting-edge AI projects with leading LLM companies. Offer Details: Commitments Required : At least 4 hours per day and minimum 20 hours per week with overlap of 4 hours with PST. Employment type : Contractor assignment (no medical/paid leave) Duration of contract : 4-8 weeks; [expected start date is next week] Evaluation Process: Delviery review of candidate profile and delivery interview for 45-60 mins
Pay
Not disclosed
Turing AI Advancement Work doesn’t publish a rate for this listing - it isn’t missing from our data.
About Turing AI Advancement Work
Turing AI Advancement Work's connection to Turing: Direct project marketplace. Coding, data science, model evaluation, voice and domain review
Ownership confidence: Confirmed. See the full network profile →
Similar open roles
Frequently asked questions
How do I apply for the Senior Software Engineer – Python (LLM Evaluation & Repository Validation) role at Turing AI Advancement Work?
Applications go through Turing AI Advancement Work's own portal - HumanSourcer links to the listing and never collects applications or handles hiring. Turing AI Advancement Work's access model is "Apply" (Apply to listed projects). This listing was first seen here on September 1, 2026 and was still live at the last check.
Is the Senior Software Engineer – Python (LLM Evaluation & Repository Validation) role at Turing AI Advancement Work remote?
Yes - Turing AI Advancement Work lists this role as remote.
What does the Senior Software Engineer – Python (LLM Evaluation & Repository Validation) role at Turing AI Advancement Work pay?
Turing AI Advancement Work doesn't state a rate on this listing. HumanSourcer never fills that gap with an estimate, so the absence is the source's, not a hole in this page. Turing AI Advancement Work's profile shows what its rate-carrying listings do advertise.
What kind of work is Senior Software Engineer – Python (LLM Evaluation & Repository Validation)?
Coding / software engineering. Coding, data science, model evaluation, voice and domain review
Is Turing AI Advancement Work a legitimate AI-training platform?
Direct project marketplace - and that relationship is rated "Confirmed" here because it can be checked against work.turing.com rather than taken on Turing AI Advancement Work's word. Confidence describes how well the ownership is evidenced. It is not a rating of how the platform treats the people who work for it, which no public source covers reliably.
Get new AI-training roles by email
One weekly digest of newly listed roles across every tracked network - pay where it's disclosed, no spam, unsubscribe anytime.
Free. See how we make money.