DSML.KZ Community Hub

The AI/ML community connecting people, knowledge, research, and opportunities across Kazakhstan and Central Asia.

Community Feed

Open profile: Andrey Lyapeikov
Andrey Lyapeikov@zergomora
New DSML profile

Data Engineer · Astana, Kazakhstan

Final-year MIPT student. Focused on high-load, distributed data infrastructure for fintech, e-commerce and telecom. Stack: Go, Java, Python, Kafka, Airflow, dbt, Snowflake, Databricks

Open profile
Read news: Open datasets for Kazakh

Open datasets for Kazakh

Community news

Asel Yermekova has compiled Awesome Kazakh Datasets—a large, carefully structured collection of open datasets for the Kazakh language. It brings together data for: • NLP and LLM: QA, RAG, instruction tuning, NER, machine translation, and other tasks • speech: ASR, TTS, speech translation, emotion recognition • computer vision, OCR, and multimodal tasks For each dataset, the collection lists its size, scale, task, release year, and download link. It also separately marks resources that were announced but could not yet be independently verified. Repositories like this make life much easier: instead of spending hours searching, you can quickly see what Kazakh-language data already exists and what you can use for your research or product. If you know of a dataset that is not yet on the list, the repository is open for contributions: 🔗 https://github.com/Allessyer/awesome-kaz-datasets Thanks to Asel for this work, and to Alen Issayev for helping develop it!

Read news
Open vacancy: AI Engineer at Til-Qazyna National Scientific and Practical Center

AI Engineer at Til-Qazyna National Scientific and Practical Center

New vacancy

Astana / Office · 600,000 to 700,000 KZT Gross per month

Til-Qazyna builds digital and AI infrastructure for the Kazakh language, including foundational language resources and open-source models. The team develops linguistic tools, custom text corpora, morphological analyzers, and fine-tuned large language models. Responsibilities: • Fine-tune, evaluate, and benchmark open-source LLMs for Kazakh using LoRA, QLoRA, and PEFT. • Develop and optimize rule-based and neural NLP components, including morphological analyzers, lemmatizers, tokenizers, and terminology parsers. • Build pipelines for text crawling, cleaning, deduplication, synthetic data generation, and dataset curation. • Package NLP models into Docker and FastAPI microservices and optimize inference with vLLM, Ollama, TensorRT-LLM, or ONNX. Requirements: • 1+ years of hands-on NLP, AI, or machine learning engineering experience. • Understanding of Transformers, attention mechanisms, embeddings, and tokenization; experience with PyTorch, Hugging Face libraries, and Scikit-learn. • Experience building preprocessing pipelines for structured and unstructured text data. • Practical knowledge of Docker, Linux, and Git. • Degree in Computer Science, Data Science, or Applied Mathematics. Optional: • Experience with LLM serving and optimization, including vLLM, TGI, llama.cpp, AWQ, or GPTQ. • Experience with RAG stacks such as LangChain, LlamaIndex, ChromaDB, Qdrant, or FAISS. • Understanding of Kazakh morphological, agglutinative, or grammatical characteristics; familiarity with DPO, RLHF, or synthetic data generation.

Open vacancy

Join DSMLKZ

Create a profile, follow the relevant channels, and take part in technical community activity.