بازگشت به فرصت‌ها
لوگوی گروه حقوق آسیا | Asia Law Group
English

Data Engineer (مشهد)

گروه حقوق آسیا | Asia Law Group·خراسان رضوی·۱ هفته پیش

حضوریمیان‌سطحقراردادیتوافقی

گروه حقوق آسیا | Asia Law Group

شرح موقعیت

We are looking for a data engineer to build reliable pipelines that transform diverse Persian-language documents into accurate, traceable, and up-to-date knowledge-base content. Project details will be shared during later recruitment stages under a confidentiality agreement. We prefer on-site collaboration in Mashhad. Required Skills Preferred Qualifications A demonstrable RAG knowledge-base ingestion project is a strong advantage. Expected Deliverables and Role Scope Reliable, monitored ingestion and update pipelines; structured, traceable data; quality reports; and maintenance documentation. This role owns data infrastructure and preparation and collaborates with the AI/RAG engineer on retrieval integration. Subject-matter specialists are responsible for validating source content. Working Arrangement and Application On-site collaboration in Mashhad is preferred. Please include your current location and on-site availability. Where possible, submit a relevant project example describing your contribution, data scale, error handling, and quality controls. For confidential projects, a high-level description without disclosing protected information is sufficient.

مسئولیت‌ها

  • Build ingestion pipelines for files, APIs, and authorized web sources, including PDF, scanned documents, Word, and HTML.
  • Implement and improve Persian OCR, flagging low-quality outputs for review.
  • Clean and normalize Persian text while preserving headings, tables, footnotes, page references, and relationships between sections.
  • Design storage for original files, processed text, and metadata, including identifiers, topics, versions, permissions, publication dates, validity periods, and ingestion timestamps.
  • Detect duplicates, preserve meaningful version differences, and maintain historical records and data lineage.
  • Structure narrative records and case studies into problems, actions, outcomes, and linked evidence; remove identifying information before shared use.
  • Prepare, chunk, and index data according to designs developed with the AI/RAG engineer.
  • Monitor sources, detect changes, perform incremental updates, and propagate modifications and deletions to processed text, chunks, search indexes, and vector data.
  • Build scheduled and batch workflows with retries, failure recovery, reprocessing, and duplicate-safe execution.
  • Implement automated quality checks, exception review workflows, logging, and alerts for processing failures or delayed updates.
  • Optimize processing speed, cost, and resource usage for large document collections.
  • Enforce access controls, user data isolation, and retention and deletion policies.
  • Document data structures, workflows, and backup and recovery procedures.
  • Strong Python and SQL skills and practical experience with ETL/ELT pipelines.
  • Experience with relational databases such as PostgreSQL and file or object storage.
  • Experience processing documents, using OCR, and handling Persian text challenges.
  • Familiarity with API integration and authorized web data extraction.
  • Ability to design data schemas, manage metadata, deduplicate records, and maintain versions.
  • Experience with workflow scheduling, error handling, and data quality controls.
  • Proficiency with Git and familiarity with Linux, Docker, and maintainable code development.
  • Ability to read English technical documentation and collaborate with AI, backend, and content teams.
  • Careful handling of confidential information and adherence to data protection requirements.
  • Understanding of LLMs and the differences between model training, fine-tuning, and RAG.
  • Experience building document ingestion pipelines for AI knowledge bases while preserving structure, provenance, and versions.
  • Familiarity with chunking, embeddings, vector indexing, keyword and semantic search, and hybrid retrieval.
  • Experience maintaining knowledge-base updates and deletions across dependent components.
  • Familiarity with preparing training, fine-tuning, and evaluation datasets, including formatting, deduplication, and preventing train–test leakage.
  • Ability to troubleshoot data, chunking, and metadata issues with the AI/RAG engineer.
  • Experience with Qdrant, pgvector, Elasticsearch, or OpenSearch.
  • Experience with Airflow, Prefect, Dagster, or similar orchestration tools.
  • Experience extracting complex Persian document layouts, tables, and footnotes.
  • Experience with parallel processing, task queues, large document collections, change detection, and data lineage.
  • Experience detecting personal information and preparing confidential data.

نیازمندی‌ها

  • SQL
  • Python
  • Data engineer
  • Big Data