About the Role
We are looking for a data engineer to build reliable pipelines that transform diverse Persian-language documents into accurate, traceable, and up-to-date knowledge-base content.
Project details will be shared during later recruitment stages under a confidentiality agreement. We prefer on-site collaboration in Mashhad.
Responsibilities
• Build ingestion pipelines for files, APIs, and authorized web sources, including PDF, scanned documents, Word, and HTML.
• Implement and improve Persian OCR, flagging low-quality outputs for review.
• Clean and normalize Persian text while preserving headings, tables, footnotes, page references, and relationships between sections.
• Design storage for original files, processed text, and metadata, including identifiers, topics, versions, permissions, publication dates, validity periods, and ingestion timestamps.
• Detect duplicates, preserve meaningful version differences, and maintain historical records and data lineage.
• Structure narrative records and case studies into problems, actions, outcomes, and linked evidence; remove identifying information before shared use.
• Prepare, chunk, and index data according to designs developed with the AI/RAG engineer.
• Monitor sources, detect changes, perform incremental updates, and propagate modifications and deletions to processed text, chunks, search indexes, and vector data.
• Build scheduled and batch workflows with retries, failure recovery, reprocessing, and duplicate-safe execution.
• Implement automated quality checks, exception review workflows, logging, and alerts for processing failures or delayed updates.
• Optimize processing speed, cost, and resource usage for large document collections.
• Enforce access controls, user data isolation, and retention and deletion policies.
• Document data structures, workflows, and backup and recovery procedures.
Required Skills
• Strong Python and SQL skills and practical experience with ETL/ELT pipelines.
• Experience with relational databases such as PostgreSQL and file or object storage.
• Experience processing documents, using OCR, and handling Persian text challenges.
• Familiarity with API integration and authorized web data extraction.
• Ability to design data schemas, manage metadata, deduplicate records, and maintain versions.
• Experience with workflow scheduling, error handling, and data quality controls.
• Proficiency with Git and familiarity with Linux, Docker, and maintainable code development.
• Ability to read English technical documentation and collaborate with AI, backend, and content teams.
• Careful handling of confidential information and adherence to data protection requirements.
Preferred Qualifications
• Understanding of LLMs and the differences between model training, fine-tuning, and RAG.
• Experience building document ingestion pipelines for AI knowledge bases while preserving structure, provenance, and versions.
• Familiarity with chunking, embeddings, vector indexing, keyword and semantic search, and hybrid retrieval.
• Experience maintaining knowledge-base updates and deletions across dependent components.
• Familiarity with preparing training, fine-tuning, and evaluation datasets, including formatting, deduplication, and preventing train–test leakage.
• Ability to troubleshoot data, chunking, and metadata issues with the AI/RAG engineer.
• Experience with Qdrant, pgvector, Elasticsearch, or OpenSearch.
• Experience with Airflow, Prefect, Dagster, or similar orchestration tools.
• Experience extracting complex Persian document layouts, tables, and footnotes.
• Experience with parallel processing, task queues, large document collections, change detection, and data lineage.
• Experience detecting personal information and preparing confidential data.
A demonstrable RAG knowledge-base ingestion project is a strong advantage.
Expected Deliverables and Role Scope
Reliable, monitored ingestion and update pipelines; structured, traceable data; quality reports; and maintenance documentation.
This role owns data infrastructure and preparation and collaborates with the AI/RAG engineer on retrieval integration. Subject-matter specialists are responsible for validating source content.
Working Arrangement and Application
On-site collaboration in Mashhad is preferred. Please include your current location and on-site availability.
Where possible, submit a relevant project example describing your contribution, data scale, error handling, and quality controls. For confidential projects, a high-level description without disclosing protected information is sufficient.