Back to careers

Fine-tuning, reward models and agent learning

Research Engineer — Agent Learning & Training

The mission: Improve agents and specialist models using verified outcomes, carefully designed training data and measurable gains on independent evaluations.

We hire according to current project needs. We welcome expressions of interest in these roles; timing and scope depend on our priorities.

Where you’d apply it today

Vowdo

Explore models for enterprise workflow decisions, grading and verification, using approved feedback while preserving execution permissions and human review.

Choose

Explore specialist models for evidence assessment and research quality where they outperform an existing-model baseline on independent evaluations.

What you’ll do

  • Build training and preference datasets from consented, sanitized workflow traces and human-verified outcomes; document provenance and keep evaluation sets independent.
  • Establish existing-model baselines, then develop specialist models through supervised fine-tuning, parameter-efficient adaptation and distillation where the evidence supports it.
  • Investigate reward models, preference optimization and reinforcement learning for agent behavior when the task has reliable feedback and a controlled evaluation environment.
  • Compare training methods with reproducible experiments and held-out workflow tests; assess task success, generalization, safety regressions, inference latency and cost.
  • Build training and inference pipelines with versioned data, checkpoints and rollback paths; deploy model changes only after review and regression checks.
  • Work with evaluation and systems engineers to turn verified feedback into reviewed improvements. Keep model learning separate from live execution and never turn unverified agent self-assessments into training targets.

What you’ll bring

  • Strong Python and PyTorch skills, with hands-on model training, debugging and reproducible experiment management.
  • Practical experience fine-tuning or distilling language models, preparing datasets and measuring results against a baseline.
  • Understanding of optimization, overfitting, data contamination and generalization; design independent tests for a claimed improvement.
  • Ability to build and profile model inference pipelines, reason about GPU memory and choose useful trade-offs between model quality, latency and cost.
  • Good judgment about consent, sensitive data, verified feedback and the limits of automated grading.

Helpful experience

  • Reward modeling, preference optimization such as DPO, or reinforcement learning with verifiable rewards.
  • Agent training environments, trajectory datasets, process supervision or learning from human corrections.
  • Distributed training, parameter-efficient adaptation, quantization or serving specialist models within constrained compute budgets.

Why you’ll find it rewarding

  • Work on the link between reliable feedback and measurable improvements in useful agent behavior.
  • Choose the simplest effective training approach and validate it against real workflow requirements.
  • Develop reusable model capabilities with evaluation and systems engineers, from experiment to reviewed release.

How to apply

Send a short introduction and links to relevant code, projects or a portfolio; a CV is optional. Tell us what you built, a failure you investigated, and how you checked whether it worked. We value demonstrated ability over particular degrees or years of experience.

Apply by emailOr write to us directly: contact@dpintelli.com

Opens your email app with a prepared draft addressed to us.