Job Title: AI Application & Prompt Specialist
Required experience (3–7 years):
3+ years building software systems with hands-on Python development (production code, packaging, testing)
Practical experience integrating LLMs and building RAG pipelines; experience with LangChain and/or LangGraph or equivalent agent frameworks
Experience with vector databases and embeddings workflows (Qdrant, OpenSearch, FAISS, Milvus or similar)
Cloud deployment experience (AWS required or strongly preferred; experience with Bedrock or other cloud model-hosting services advantageous)
Containerization (Docker), CI/CD pipelines, and basic infra-as-code or orchestration knowledge
Familiarity with REST/async APIs, WebSockets, and relational DBs (PostgreSQL or equivalent)
Demonstrable prompt engineering skills and experience measuring prompt performance and reducing hallucinations
Strong engineering practices: modular code, unit tests, code reviews, documentation, and reproducible deployments
Clear communication and cross-functional collaboration skills
Roles and Responsibilities:
Design, build and productionize internal AI solutions that combine strong prompting/LLM experience with production-grade Python engineering, agent orchestration, and secure enterprise integrations. Deliver reusable frameworks, observability, and guardrails so business teams can adopt LLM-enabled features safely and reliably.Build reusable SDKs, libraries and microservices for LLM integrations (prompt templates, prompt chaining, tool calling, caching, retry & backoff logic)
Architect and implement multi-agent orchestration and workflow systems using LangGraph and/or LangChain patterns to support supervisor/worker and tool-calling agent designs
Implement secure RAG pipelines: document ingestion, chunking strategies, embeddings, vector DB integration (Qdrant, OpenSearch, FAISS-style stores) and retrieval tuning
Write production Python code (APIs, async messaging, WebSockets, background workers) to integrate LLMs (cloud-hosted or private) and vector stores; package services with containers and CI/CD
Deploy and operate LLM components on cloud platforms (AWS + Bedrock or equivalent; familiarity with Azure OpenAI is a plus) and manage secrets/OAuth2 and RBAC integrations
Define and run systematic prompt evaluation and monitoring (accuracy, hallucination, cost-per-response); implement guardrails and automated tests for prompt suites (unit tests, A/B experiments)
Integrate AI features with enterprise systems (ServiceNow, SAP SuccessFactors, internal HR/ERP systems) to enable end-to-end workflows (ticket creation, approvals, data updates)
Implement observability/alerting (latency, throughput, cost, drift, hallucination rates) and incident handling; produce runbooks and monitoring dashboards
Collaborate with product, legal, security and domain experts to ensure privacy, compliance and acceptable use; implement technical controls (input redaction, auditing, access controls)
Document patterns, run enablement sessions and deliver onboarding materials for internal developers and business users
Primary Skills:
Python (production-grade code, packaging, typing)
LangChain and LangGraph (agent orchestration, prompt chaining, tool calling)
RAG: ingestion, chunking strategies, embeddings, vector search and retrieval tuning
Cloud: AWS (S3, Lambda, ECS/EKS, Bedrock or model-hosting services); familiarity with Azure OpenAI helpful
Databases: PostgreSQL, session state and audit logging best practices
DevOps: Docker, CI/CD, monitoring/observability tools, basic Kubernetes concepts
Security: OAuth2, role-based access control, input filtering/redaction, audit trails
Prompting & evaluation capabilities:
Craft and iterate high-quality prompts and templates for retrieval, summarization, instruction following and tool use
Build reusable prompt libraries and parameterized templates for scale
Implement automated prompt testing and evaluation pipelines (unit tests, A/B experiments, quantitative metrics)
Define and monitor metrics: accuracy, usefulness, hallucination rate, latency, and cost per useful response
Design mitigation patterns for failure modes (refusal strategies, retrieval confidence thresholds, fallback flows)
Governance, compliance & security
Apply enterprise AI usage rules: use approved/private model instances, avoid uploading personal or strictly confidential data into external models, and enforce output review processes where required
Implement technical enforcement: input filtering/redaction, access controls, logging/auditing and support approval workflows for publishing outputs
Collaborate with information owners, legal and security teams to ensure policies are enforced and audits supported
Secondary Skills:
Hands-on experience with LangSmith or similar evaluation/orchestration tooling
Experience with LLM fine-tuning or instruction-tuning workflows
Experience building internal SDKs, admin panels or developer platforms for AI adoption
Domain knowledge in industrial automation, manufacturing, connectivity or related fields
Interview & assessment suggestions:
Take-home task: build a small Python microservice that performs RAG using a vector DB, exposes a prompt template interface, includes unit tests and basic monitoring metrics
Live exercise: iterate prompts for a defined internal workflow, demonstrate evaluation choices and explain failure modes and mitigations
System design: outline architecture for safe LLM integration at enterprise scale (ingestion, retrieval, orchestration, monitoring, access control)
Request GitHub or code samples, architecture diagrams, and runbook/monitoring artifacts where available
Professional Competencies:
- Technical Expertise
- Analytical Skills
- Demonstrate Ownership
- Communication
- Leadership
- Execution
- Technological Insight
- Process Optimization Knowledge
- Business & Financial Acumen
- Trust and Collaborate
Culture Competencies:
- Become better everyday
- Break new ground
- Champion customers satisfaction
- Demonstrate ownership
- Trust and collaborate