Location: On-site | Employment Type: Full-time | Experience: 24 years
About ZharfaTech
At ZharfaTech, we build the future of Generative AI from Agentic and Conversational AI to Speech and Image technologies. Our team tackles cutting-edge challenges using the latest tools and frameworks, with a mission to push the boundaries of artificial intelligence and deliver scalable, high-performance solutions.
About the Role
As a Mid-level AI Engineer, you will build, fine-tune, and optimize production-grade speech and audio ML systems. You will work hands-on with PyTorch and CUDA to train, adapt, and serve models across the speech-processing stack automatic speech recognition (ASR), speaker diarization, voice activity detection (VAD), speech enhancement/separation, and forced alignment and expose them through reliable, observable service APIs. We are looking for an engineer who is strong in Python and modern ML frameworks, has practical experience fine-tuning sequence models (especially ASR), and is eager to grow deeper in speech/audio and GPU systems while shipping real product impact.
What You'll Do
Model Fine-tuning & Adaptation: Fine-tune and adapt speech and audio models with a strong focus on ASR (wav2vec2 / CTC, Whisper, encoder-decoder and LLM-decoder hybrids) on domain and in-language data using techniques such as LoRA/QLoRA, PEFT, and CTC/decoder fine-tuning to improve accuracy on low-resource and code-switched settings.
Data & Evaluation: Curate, augment, and clean training corpora (transcripts, audio normalization, segmentation); design WER/CER and speaker-attributed evaluation pipelines, run error analysis, and track regressions across model versions.
Pipeline Development: Build and maintain stages of a multi-stage speech pipeline ASR, VAD, diarization, enhancement/separation, forced alignment, and LLM-based transcript correction using PyTorch, Hugging Face Transformers, and related model libraries.
Inference & Serving: Integrate and run pretrained and fine-tuned models on GPU, implement ensemble fusion, post-processing, and word-level alignment, and serve them through high-performance APIs with real-time progress streaming.
GPU & Memory Management: Profile and tune VRAM usage, batch sizing, and execution graphs; manage multi-model lifecycles with controlled loading/eviction and per-device placement on modern GPUs.
Observability & Reliability: Instrument services with metrics and health checks, debug end-to-end pipeline issues, handle out-of-memory recovery, and improve throughput and latency under concurrent workloads.
Collaboration: Partner with backend and frontend engineers and product teams, and contribute to code reviews, testing, and technical documentation.
What You'll Bring
Experience: 24 years of hands-on experience in Python ML or software engineering, with meaningful exposure to model training, fine-tuning, or inference.
ASR Fine-tuning: Demonstrable experience fine-tuning automatic speech recognition models (e.g., wav2vec2/CTC, Whisper, or similar) on real datasets, including data preparation, training runs, and quantitative evaluation (WER/CER).
ML Frameworks: Solid proficiency with PyTorch and the Hugging Face ecosystem (transformers, datasets, accelerate); comfort loading, fine-tuning, running, and debugging transformer/audio models on GPU.
Backend Engineering: Strong async Python skills and experience building REST APIs with FastAPI or a comparable framework; understanding of asyncio, concurrency, and thread/executor patterns.
GPU/CUDA: Familiarity with running models on CUDA GPUs device placement, VRAM awareness, mixed-precision training, and troubleshooting OOM errors.
Audio Fundamentals: Understanding of digital audio (sample rate, waveforms/spectrograms, PCM, format transcoding) and libraries such as librosa, soundfile, or torchaudio.
Tooling: Experience with Git, Docker, and dependency management; comfort with Linux and command-line workflows.
Problem-Solving & Communication: A debug-it-to-the-root mindset with the ability to read research code and vendor libraries, plus good written and verbal communication skills.
Nice to Haves
Speech/Audio Depth: Experience with speaker diarization, VAD, speech enhancement/separation, forced alignment, or sound event detection.
Low-resource & Multilingual ASR: Experience adapting ASR for low-resource or morphologically rich languages, cross-lingual transfer, or language model rescoring/shallow fusion.
Parameter-Efficient Training: Hands-on experience with LoRA/QLoRA, PEFT, or distributed/efficient fine-tuning of large speech models.
Model Optimization: Familiarity with quantization (FP8/AutoAWQ), torch.compile, or inference-serving frameworks (vLLM, TensorRT-LLM, ctranslate2).
LLM & RAG: Experience with LLM inference, prompt engineering, or retrieval-augmented generation (RAG) using vector databases.
MLOps: Exposure to MLflow, LangFuse, or similar tools for experiment tracking, dataset versioning, and observability.
Open Source: Contributions to open-source AI/speech projects or published work in the speech/audio domain.
Why Join ZharfaTech?
Real Impact: Own and ship core parts of production speech-AI systems used by real users.
Grow Fast: Work alongside senior engineers on cutting-edge speech, audio, and LLM systems with real mentorship.
Modern Stack: PyTorch, Hugging Face, FastAPI, vLLM, and the latest open speech models on modern GPU hardware.
Innovation Culture: Experiment with emerging tools and frameworks, and turn research into shipped product.
Mission-Driven: Help build AI that makes spoken language accessible, searchable, and actionable.