← Back to jobs

Solution Architect – AI Voice Backend

Skills

aiaksandroidapiazureazure devopschatgptci cdcloudconfluencedevopsdistributed systemsembeddingsfastapigdprgithubgrpcinfrastructure as codekuberneteslangchainlanggraphllmmicroservicesmongodb

Description

Project description

Our client is advancing its in-vehicle voice assistant into an intelligent, AI-powered companion. Large-language-model capabilities (Azure OpenAI / ChatGPT) have been running in production across vehicles with new E³-architecture models featuring enhanced voice functions from the factory. The customer backend is the cloud AI orchestration service behind this: it receives requests from the vehicle, classifies and routes them, orchestrates the LLM, tool services and specialised agents, and returns an answer or action to the car. DXC Luxoft serves as the end-to-end delivery partner, working in a joint product team with the client's engineers on the Azure platform (AKS, Azure OpenAI, AI Foundry, Managed Identity, Azure Monitor, Azure DevOps). This role owns the architecture of the backend service in series operation — a system that is already live across a large vehicle fleet and must stay available and backward-compatible while significant new capability is added: streaming across the full ASR → LLM → TTS chain, multi-intent handling, barge-in, a guardrails layer for deterministic vehicle-safe answers, agent routing, and new tool integrations. It also owns the vehicle ↔ backend interface contract: how the car talks to backend, across several vehicle generations at once.

Responsibilities

  • Own and evolve the system architecture of the backend service: routing, business logic, orchestration, service decomposition, and the structural discipline (hexagonal architecture / ports and adapters) that keeps the platform changeable over a multi-year lifecycle.
  • Define and govern the interface concept between vehicle, LLM, tool services and agents — the formal contract layer, its versioning strategy, and its evolution across vehicle generations.
  • Architect the vehicle-facing integration: REST and gRPC/protobuf APIs, Viwi, AIDL toward the Android Automotive side, OAuth2 and mTLS, client IdP integration, and the client's internal host and telemetry interfaces.
  • Design the streaming architecture — migrating AI service calls to streaming for responsiveness: incremental chat completion, real-time ASR transcription, TTS playback during synthesis; with buffering, connection management, error handling, and clean cancellation/interruption semantics (the foundation for barge-in).
  • Own the LLM integration layer: prompt normalisation and pre-/post-processing, response format control for vehicle display, context preparation and dialogue management, and deterministic vehicle answers via system prompts and a guardrails engine — in a domain where a wrong answer is a safety and brand issue, not just a quality one.
  • Design the agent orchestration mechanism that decides whether a request is answered by the LLM, an internal tool, or an external agent — including multi-intent handling and the routing model behind it (LangGraph).
  • Architect tool and agent integrations: navigation with semantic location resolution, media/entertainment search, knowledge queries, calendar and mail, POI and places services, and vehicle data services.
  • Define the RAG and persistence architecture: embedding strategy and lifecycle, vector store design (PostgreSQL/pgvector), schema design, tuning and migrations, and the division of responsibility between relational, document and vector stores.
  • Produce scaling and load concepts for series operation, and the security, data-protection and compliance concept per automotive standards — including data minimisation in telemetry and traces.
  • Own the observability architecture end to end: OpenTelemetry distributed tracing, LangFuse for prompt tracing and evaluation, Azure Monitor metrics and dashboards, alert thresholds and routing — so latency can be attributed per component and failure chains reconstructed.
  • Make the architecture hold its operational targets: high availability in the 99.5% SLO range, incident resolution inside 24 hours, business-hours 2nd/3rd-level support, and structured version, release and deployment management.
  • Guarantee backward compatibility for existing vehicle generations — vehicles from model year 2021 onward must keep working as new features ship.
  • Set and enforce engineering standards with the development teams: API and interface review, Python toolchain and code quality gates, CI/CD pipeline design on Azure DevOps (including self-hosted runners) and GitHub, IaC module structure in Terraform/Terragrunt.
  • Act as technical counterpart to the client's architects and to the vehicle, backend and UX teams; maintain architecture documentation in Confluence and work within the client's requirements management tooling; provide technical direction to the distributed development and AI Ops team in Scrum.

SKILLS

Must have

  • 8+ years in backend/distributed systems engineering, including 3+ years as architect or technical lead with end-to-end ownership of a production system's design.
  • Expert Python: FastAPI, async, and hands-on experience with streaming architectures (SSE/WebSocket/gRPC streaming, back pressure, cancellation).
  • Proven experience architecting LLM-based production systems — orchestration of tool- and agent-based workflows with LangChain / LangGraph or equivalent, prompt strategies, structured outputs, LLM constraint handling and guard-railing.
  • Azure OpenAI Service or OpenAI API in production, including model version migration and the compatibility problems it creates.
  • Strong API and interface design: REST and gRPC/protobuf, versioning and backward compatibility strategies, OAuth2 and mTLS for secure service-to-service communication.
  • PostgreSQL at depth: schema design, tuning, migrations (Alembic); plus working experience with embeddings and vector search for RAG (pgvector or equivalent).
  • Kubernetes / container orchestration (deployment, scaling, secrets/config, networking, load balancing) and Azure cloud architecture (AKS or Container Apps, Managed Identity, Key Vault, Monitor).
  • Infrastructure as Code with Terraform: reusable modules, environment separation, remote state and locking, CI/CD integration, lifecycle management.
  • Demonstrable responsibility for running a production service: monitoring, alerting, distributed tracing with OpenTelemetry, incident and problem management, patch and release management, latency and cost optimization against defined targets.

Nice to have

• Automotive or in-vehicle software background: MIB3 / E³ / SDV platforms, Android Automotive (AAOS) and AIDL interfaces, Viwi protocol, vehicle telemetry and configuration services, SOA / microservice / zonal E/E architectures. • Experience with LLM evaluation and prompt regression practice: LangFuse, DeepEval, Ragas, PromptFlow; metrics such as faithfulness, hallucination rate, latency P95, WER/CER. • ASR/TTS service integration at the streaming level (Azure Speech, Whisper, Cerence, Google TTS) and awareness of in-cabin speech conditions (noise, far-field microphones, multi-speaker). • Emotion/sentiment APIs (HumeAI) and empathic dialogue design. • Broader persistence experience: MongoDB (pymongo/beanie), CosmosDB, and a clear opinion on when each belongs in the picture. • Semantic caching and token-cost optimization at fleet scale. • Awareness of automotive quality, security and AI governance frameworks: A-SPICE, ISO/SAE 21434, TISAX, GDPR/privacy-by-design, EU AI Act implications. • Experience taking over and stabilising a running system from a previous supplier (knowledge transfer, reverse-documenting an inherited codebase)

Get similar jobs in Germany by email

We'll email you when new jobs similar to this one appear.

Similar jobs

Explore more Solutions Architect jobs in Germany.

No similar openings right now. Refine your search or create an alert above.