TekionFound on the web

Posted 2w ago

Staff Software Development Test Engineer - AI Evaluation

Pay

Not stated

Location

Bangalore HQ

On site

Skills this role screens for

pythoncloudai/mldata scienceanalyticsproduct managementall cost centers

Opens on Tekion's site. Sign in first and we keep it in your pipeline.

Also

ABOUT TEKION

Positively disrupting an industry that has not seen any innovation in over 50 years, Tekion has challenged the paradigm with the first and fastest cloud-native automotive platform that includes the revolutionary Automotive Retail Cloud (ARC) for retailers, Automotive Enterprise Cloud (AEC) for manufacturers and other large automotive enterprises and Automotive Partner Cloud (APC) for technology and industry partners. Tekion connects the entire spectrum of the automotive retail ecosystem through one seamless platform. The transformative platform uses cutting-edge technology, big data, machine learning, and AI to seamlessly bring together OEMs, retailers/dealers and consumers. With its highly configurable integration and greater customer engagement capabilities, Tekion is enabling the best automotive retail experiences ever. Tekion employs close to 3,000 people across North America, Asia and Europe.

The work

About the Role

We are looking for a highly motivated Staff SDET – AI Evaluation to join Tekion’s AI Platform  team. Evaluation is the backbone of trustworthy AI: as Tekion scales from a handful of AI agents

to 100+ across Service, Sales, F&I, and Analytics, this role builds the evaluation platform and frameworks that let every ML team measure, trust, and improve the quality of AI outputs.In this role, you will be responsible for defining and building Tekion’s AI evaluation capabilities as a shared platform service. You will work closely with ML Engineers, Data Scientists, the AI Platform team, and Product Management to design evaluation datasets, automated scoring pipelines, and quality metrics that quantify the accuracy, consistency, and safety of AIgenerated outputs across the organization. You will own the systems that answer “is this model or agent good enough to ship, and is it

staying good in production?” — from offline benchmarks and LLM-as-judge pipelines to online evaluation and continuous quality monitoring. You will also use AI and LLMs to scale evaluation itself, building automated judges and synthetic datasets that expand coverage faster than manual review ever could.

The work

What You’ll Do

  • Develop a deep understanding of Tekion’s AI agents, ML models, and the quality dimensions that matter for each business domain.
  • Design, enhance, and own Tekion’s AI evaluation infrastructure as a shared capability used across ML teams.
  • Create, curate, and maintain evaluation datasets (evals) and golden/ground-truth sets across use cases and domains.
  • Define quality metrics for AI outputs — accuracy, relevance, faithfulness/groundedness, consistency, safety, and task success.
  • Build automated scoring pipelines, including LLM-as-judge, rubric-based, and referencebased evaluation methods.
  • Validate user intents and measure response accuracy and consistency for AI-powered capabilities such as the Analytics Agent.
  • Identify hallucinations, unsafe or biased outputs, and edge cases; design targeted eval suites to catch them.
  • Build both offline evaluation (pre-release benchmarking) and online evaluation (production quality monitoring, A/B, drift detection).
  • Establish evaluation gates in CI/CD so model, prompt, or data changes are quality-checkedbefore release.
  • Develop dashboards and reporting that make AI quality visible and actionable for ML and product teams.
  • Use AI/LLMs to scale evaluation — automated judges, synthetic data generation, and eval \tooling
  • Champion evaluation and responsible-AI quality best practices across the organization.

Who they want

What We’re Looking For

  • 8+ years in SDET, quality engineering, ML engineering, or data science, with hands-on experience building evaluation or measurement systems — or a strong SDET background with deep LLM/ML fluency.
  • Strong programming skills in Python, with the ability to build robust, reusable evaluation pipelines and tooling.
  • Deep understanding of ML/LLM evaluation, benchmark design, and the pitfalls of evaluating non-deterministic

Tekion

Found on the web. You apply on the company's own site

Where this listing comes from

From Tekion's own careers system (Ashby).

insiderOne never charges students to apply. Never pay anyone for an internship. Check an offer you received

Prove Python first

This role screens for Python. Sit a ten-minute live challenge, get a signed receipt, and the employer sees it verified on your application instead of claimed.

Prove it, then apply

Ask about this role

SCOUT, an AI agent, not a person, answers from this ad, and says so when the ad doesn't cover it. No account, and nothing you type here is kept.

More ai and machine learning roles

Every ai and machine learning role here

insiderOne University

Get certified for roles like this: AI Fundamentals Associate

Free, about 2 hours, ending in a project a person reviews and a certificate any employer can check. Put it on this application and every one after it.

Start the certification

Interned at Tekion? Report what it paid

From the Wire Desk