شرح موقعیت
We are looking for a data engineer to build reliable pipelines that transform diverse Persian-language documents into accurate, traceable, and up-to-date knowledge-base content. Project details will be shared during later recruitment stages under a confidentiality agreement. We prefer on-site collaboration in Mashhad. Required Skills Preferred Qualifications A demonstrable RAG knowledge-base ingestion project is a strong advantage. Expected Deliverables and Role Scope Reliable, monitored ingestion and update pipelines; structured, traceable data; quality reports; and maintenance documentation. This role owns data infrastructure and preparation and collaborates with the AI/RAG engineer on retrieval integration. Subject-matter specialists are responsible for validating source content. Working Arrangement and Application On-site collaboration in Mashhad is preferred. Please include your current location and on-site availability. Where possible, submit a relevant project example describing your contribution, data scale, error handling, and quality controls. For confidential projects, a high-level description without disclosing protected information is sufficient.
مسئولیتها
- Build ingestion pipelines for files, APIs, and authorized web sources, including PDF, scanned documents, Word, and HTML.
- Implement and improve Persian OCR, flagging low-quality outputs for review.
- Clean and normalize Persian text while preserving headings, tables, footnotes, page references, and relationships between sections.
- Design storage for original files, processed text, and metadata, including identifiers, topics, versions, permissions, publication dates, validity periods, and ingestion timestamps.
- Detect duplicates, preserve meaningful version differences, and maintain historical records and data lineage.
- Structure narrative records and case studies into problems, actions, outcomes, and linked evidence; remove identifying information before shared use.
- Prepare, chunk, and index data according to designs developed with the AI/RAG engineer.
- Monitor sources, detect changes, perform incremental updates, and propagate modifications and deletions to processed text, chunks, search indexes, and vector data.
- Build scheduled and batch workflows with retries, failure recovery, reprocessing, and duplicate-safe execution.
- Implement automated quality checks, exception review workflows, logging, and alerts for processing failures or delayed updates.
- Optimize processing speed, cost, and resource usage for large document collections.
- Enforce access controls, user data isolation, and retention and deletion policies.
- Document data structures, workflows, and backup and recovery procedures.
- Strong Python and SQL skills and practical experience with ETL/ELT pipelines.
- Experience with relational databases such as PostgreSQL and file or object storage.
- Experience processing documents, using OCR, and handling Persian text challenges.
- Familiarity with API integration and authorized web data extraction.
- Ability to design data schemas, manage metadata, deduplicate records, and maintain versions.
- Experience with workflow scheduling, error handling, and data quality controls.
- Proficiency with Git and familiarity with Linux, Docker, and maintainable code development.
- Ability to read English technical documentation and collaborate with AI, backend, and content teams.
- Careful handling of confidential information and adherence to data protection requirements.
- Understanding of LLMs and the differences between model training, fine-tuning, and RAG.
- Experience building document ingestion pipelines for AI knowledge bases while preserving structure, provenance, and versions.
- Familiarity with chunking, embeddings, vector indexing, keyword and semantic search, and hybrid retrieval.
- Experience maintaining knowledge-base updates and deletions across dependent components.
- Familiarity with preparing training, fine-tuning, and evaluation datasets, including formatting, deduplication, and preventing train–test leakage.
- Ability to troubleshoot data, chunking, and metadata issues with the AI/RAG engineer.
- Experience with Qdrant, pgvector, Elasticsearch, or OpenSearch.
- Experience with Airflow, Prefect, Dagster, or similar orchestration tools.
- Experience extracting complex Persian document layouts, tables, and footnotes.
- Experience with parallel processing, task queues, large document collections, change detection, and data lineage.
- Experience detecting personal information and preparing confidential data.
نیازمندیها
- SQL
- Python
- Data engineer
- Big Data