AI Research Engineer, Computer Vision & VLMs at Palona AI | CA | Rezi

AI Research Engineer, Computer Vision & VLMs at Palona AI

AI Research Engineer, Computer Vision & VLMs

Palona AI · CA

1 weeks ago

AI Research Engineer, Computer Vision & VLMs

Palona AI · CA

11 days ago
Resume preview

Impress employers and recruiters.
Choose from hundreds of resume examples.

Target Resume Now
Resume preview

Tailor your resume to this AI Research Engineer, Computer Vision & VLMs role.

Rezi rewrites your resume against Palona AI's job description. Free.

Resume score gauge reading 58 out of 100

Don't guess if your resume is good enough.

See how it scores against the AI Research Engineer, Computer Vision & VLMs posting at Palona AI — free, in seconds.

About the Role

Palona is developing AI for the physical world, starting with restaurants. We are seeking an AI Research Engineer with expertise in computer vision and vision-language models (VLMs) to create the visual intelligence for our products. This role involves research and development of image and video understanding, spatiotemporal reasoning, and multimodal models to derive insights and actions from visual data in restaurant settings.

Responsibilities

  • Develop computer vision and VLM approaches for scene understanding, object detection and tracking, activity recognition, and event understanding across video.
  • Adapt, fine-tune, and evaluate vision and vision-language models for visual grounding, temporal reasoning, and structured prediction.
  • Design training and adaptation strategies, including supervised fine-tuning, representation learning, distillation, and domain adaptation.
  • Build representative image and video datasets, annotation workflows, and evaluation sets that capture difficult edge cases while protecting sensitive data.
  • Create rigorous experiments and benchmarks to measure perception quality, temporal consistency, hallucinations, robustness, latency, and cost.
  • Diagnose failures caused by occlusion, lighting changes, camera placement, rare events, and domain shift, and use findings to improve data and models.
  • Partner with infrastructure and product engineers to deploy efficient inference pipelines with monitoring, quality gates, staged rollouts, and rollback paths.
  • Translate advances in computer vision, VLMs, and embodied AI into practical product capabilities, communicating evidence and tradeoffs.
  • Raise research and engineering standards through reproducible experiments, thoughtful reviews, and clear documentation.

Requirements

  • 3+ years of research or applied development experience in computer vision, multimodal learning, or a closely related field; relevant graduate research counts toward this experience.
  • Demonstrated research track record in computer vision or vision-language modeling through publications, substantial research projects, open-source contributions, or industry research.
  • Strong foundations in deep learning, visual representation learning, and experimental design, with depth in areas such as video understanding, detection and tracking, visual grounding, or multimodal reasoning.
  • Hands-on experience training, fine-tuning, or adapting computer vision models, and developing or evaluating VLMs beyond basic API integration.
  • Strong Python skills and experience with PyTorch or an equivalent deep learning framework, along with modern training and evaluation tooling.
  • Experience building datasets, designing reliable evaluations, analyzing model failures, and using ablations to understand improvements.
  • Strong software engineering judgment and the ability to turn research code into reproducible, tested systems.
  • Ability to connect modeling choices to product constraints including latency, cost, privacy, reliability, and user experience.
  • Comfort working through ambiguity and collaborating across research, engineering, and product.

Skills

  • Computer Vision
  • Vision-Language Models (VLMs)
  • Deep Learning
  • Visual Representation Learning
  • Experimental Design
  • Video Understanding
  • Detection and Tracking
  • Visual Grounding
  • Multimodal Reasoning
  • Python
  • PyTorch

Location

  • U.S.-based roles

Work Type

  • Full-time

Experience Level

  • 3+ years of research or applied development experience

Education Level

  • PhD or research-focused master’s degree in computer vision, machine learning, robotics, or a related field, or equivalent research experience.

Salary/Compensations

  • Competitive salary

Benefits

  • Stock option plan
  • Company-sponsored green card applications for strong candidates hired into U.S.-based roles, subject to eligibility
  • Medical, dental, vision, and retirement benefits
  • Family leave and short-term and long-term disability benefits
  • Paid time off and company holidays
  • Learning and development support

About the Company

  • Palona is building AI for the physical world, starting with restaurants.