FyerX

FyerX

·2 days ago

Nlp data scientist

Apply now

Location

remote (anywhere)

Commitment

Full Time

Level

Middle (2-4 years)

Required skills

Natural Language ProcessingLarge Language Model Fine-TuningPythonPyTorchHugging Face TransformersParameter-Efficient Fine-TuningLoRAQLoRAReinforcement Learning from Human FeedbackDirect Preference OptimizationDataset PreparationModel QuantizationDistributed Deep LearningModel EvaluationCUDAMLOps

Job Description

Job Details

Employment Type: Contract

Work Mode: Remote

Location: Offshore

Total Experience Required: 4 to 8 years

Relevant Experience Required: 3+ years of dedicated natural language processing (NLP) and hands-on Large Language Model (LLM) fine-tuning experience

Mandatory Certification: Google Cloud Certified Professional Machine Learning Engineer or AWS Certified Machine Learning - Specialty

Job Summary

We are seeking an experienced NLP Data Scientist / LLM Fine-Tuning Specialist to take ownership of our specialized open-source model optimization tracks. The ideal candidate will possess deep expertise in deep learning, dataset preparation, and parameter-efficient training methodologies to fine-tune foundational models for industry-specific terminology, domain-specific reasoning, and custom task execution.

Key Responsibilities

  • Lead LLM fine-tuning initiatives, leveraging Parameter-Efficient Fine-Tuning techniques (PEFT) including LoRA, QLoRA, Prefix Tuning, and Prompt Tuning to optimize open-source architectures (e.g., Llama, Mistral).
  • Curate, clean, and structure high-quality training datasets, implementing automated data deduplication, tokenization schemes, synthetic data generation pipelines, and human-in-the-loop validation frameworks.
  • Implement advanced reinforcement learning alignment layers, configuring Reinforcement Learning from Human Feedback (RLHF) or Direct Preference Optimization (DPO) to enforce model safety, helpfulness, and tone guardrails.
  • Optimize model footprint constraints and memory overhead, applying post-training quantization techniques (e.g., GGUF, AWQ, GPTQ) to minimize parameter degradation and compute budgets.
  • Design rigorous evaluation benchmarks and metrics panels, executing automated validation tests (e.g., BLEU, ROUGE, custom verification matrices) to audit model hallucinations, factual accuracy, and domain alignment.
  • Manage distributed deep learning training jobs, scaling pipeline configurations, tensor parallelism parameters, and gradient checkpointing scripts across multi-GPU compute blocks.
  • Collaborate with MLOps infrastructure teams, formatting completed model weight checkpoints cleanly for scalable cloud deployment and real-time inference serving layers.

Requirements

4 to 8 years of core data science or advanced machine learning engineering experience, with 3+ dedicated years actively training, evaluation, and fine-tuning natural language processing systems.

Expert-level technical mastery of Python, deep learning frameworks (PyTorch), transformer architectures (Hugging Face Transformers, Accelerate, PEFT), and vector calculations.

Deep structural understanding of attention mechanisms, tokenization constraints, context window degradation behaviors, loss function optimization, and hardware compute limitations (CUDA).

Mandatory certification: Professional ML Engineer or Specialty Machine Learning credential from a major cloud vendor (AWS/GCP).

Preferred Qualifications

Master’s or Ph.D. in Computer Science, Data Science, Computational Linguistics, or an adjacent quantitative field with a research focus on neural network text models.

Prior experience implementing custom embedding model structures or optimizing domain-specific classification layers inside constrained enterprise runtimes.

Ready to join the team?

Apply now

Similar Jobs: