به گستران | Beh Gostaran
شرح موقعیت
English: C1 – Advanced / Fluent Professional Proficiency Advanced/Fluent professional proficiency, suitable for technical discussions, meetings, documentation, and collaboration with international teams. Weak English communication may result in rejection because the role requires frequent cross-functional teamwork. The LLM / GenAI Engineer builds production AI systems that combine foundation models, retrieval, structured data, and reliable software services. The role covers RAG applications, tool-using agents, prompt and model optimization, and evaluation workflows that turn language models into dependable product capabilities You will work with applied scientists, platform engineers, and product teams to move GenAI systems from prototype to production. The work requires equal attention to answer quality, latency, cost, security, observability, and failure recovery across cloud-based deployments. Key Responsibilities What We Are Looking For
مسئولیتها
- Design and deploy RAG pipelines using Python, LangChain, LlamaIndex, or custom orchestration frameworks, integrating document ingestion, chunking, embeddings, retrieval, reranking, and response generation
- Build agentic workflows that safely connect LLMs to internal APIs, databases, search systems, and business tools with structured outputs, permissions, and recovery paths
- Develop evaluation systems using benchmark datasets, golden responses, LLM-as-judge methods, human review, and automated regression testing to measure groundedness, relevance, latency, and cost
- Fine-tune and optimize models using supervised fine-tuning, LoRA or QLoRA, prompt optimization, distillation, and model-routing strategies where appropriate
- Implement production services with Python, FastAPI, Docker, and cloud infrastructure; integrate model providers such as OpenAI, Anthropic, Google, or open-source models hosted on managed platforms
- Instrument applications with tracing, metrics, and alerts to monitor token usage, model quality, hallucination rates, latency, failures, and data drift
- Partner with security and platform teams to establish controls for PII handling, prompt injection, data access, model versioning, and safe deployment rollbacks
- 3–8 years of software engineering, machine learning engineering, or applied AI experience, including at least 1 year delivering LLM or GenAI systems to production
- Strong Python skills and experience building reliable backend services, asynchronous workflows, REST APIs, automated tests, and CI/CD pipelines
- Hands-on experience with RAG architectures, embedding models, vector databases such as Pinecone, Weaviate, Milvus, Chroma, or pgvector, and hybrid retrieval techniques
- Practical knowledge of LLM evaluation, prompt engineering, function calling, structured generation, fine-tuning, and model tradeoffs involving quality, latency, and cost
- Experience with at least one major cloud platform—AWS, GCP, or Azure—and containerized deployment using Docker and Kubernetes or an equivalent platform
- Bachelor’s or master’s degree in computer science, machine learning, engineering, mathematics, or a related technical field, or equivalent professional experience
نیازمندیها
- هوش مصنوعی
- Ai
- Python
- RESTful APIs