lingkarsolution.com

View all job listings

AI Technical Expert (LLM / Vision / Real-time Streaming)

Pusat

Contract

Hybrid

Job description

We are hiring an AI Technical Expert (LLM / Computer Vision / Real-time Streaming) to deliver production-grade AI capabilities for B2G AI and dashboard insights products. You will lead implementation from solution design to deployment and operations in regulated environments (data residency, auditability, controlled rollout), working closely with engineering, DevOps/SRE, and security teams.


Responsibilities

  • Install, deploy, and operate LLM inference on GPU servers (on-prem/private cloud), including performance tuning, scaling strategy, monitoring, and rollback/versioning.
  • Deploy and operate model serving using vLLM (and other serving stacks where appropriate), exposing secure APIs (REST/gRPC) with authentication, rate limits, audit logs, and observability.
  • Support heterogeneous infrastructure: deploy LLM workloads across NVIDIA GPU (x64) and Huawei/Ascend GPU/NPU (ARM64) environments, including containerization, drivers/runtime, and compatible serving/inference stacks.
  • Implement RAG / retrieval pipelines (ingestion, chunking, embeddings, vector search) and evaluation gates for quality and safety.
  • Perform fine-tuning/adaptation when required (LoRA/QLoRA preferred): dataset preparation (incl. de-identification), experiment tracking, and acceptance reporting.
  • Build real-time voice pipelines: live audio ingest, streaming ASR, optional diarization, and TTS integration with low latency.
  • Build real-time video streaming + computer vision: ingest (RTSP/WebRTC/SRT/HLS as needed), transcoding, GPU inference (detection/tracking/event detection), and structured event outputs for dashboards.
  • Deliver production readiness: runbooks, incident playbooks, DR/backup validation, security hardening, and go-live support.


Job requirements

5+ years in applied AI/ML engineering or ML platform engineering with production deployments.

Proven experience running LLM inference with vLLM and operating GPU-backed services reliably.

Experience deploying across x64 and ARM64, and on heterogeneous accelerators (NVIDIA + Huawei/Ascend).

Strong computer vision + real-time inference experience; solid grasp of streaming fundamentals (FFmpeg/GStreamer, low-latency patterns).

Comfortable working in B2G/regulatory constraints and stakeholder-driven delivery.

Benefits

Includes a project completion fee payable upon successful delivery and formal acceptance of agreed milestones.

Job information

Education

Bachelor Degree (S1)

Experience level

Mid Senior Level

Minimum experience

5 years

Gender

No Qualification

Published date

07 Mar 2026

Powered by

Mekari Talenta