
Evolving with Our Language, Empowered by Our Data
Türkiye's Large Language Model.
About BİLGE
What Is Bilge?
Türkiye's Domestic and National Large Language Model
BİLGE is a large language model family developed by TÜBİTAK BİLGEM, based on the unique structure of the Turkish language and our cultural heritage. It focuses on understanding the linguistic structure, semantic layers, and cultural context of the Turkish language. Rather than a “default culture,” it centers on Türkiye.
It is a product of Türkiye’s expertise and engineering capabilities in the field of artificial intelligence. With a model family ranging from 1 billion to 122 billion parameters, it caters to every scale.
A family of local language models developed for Turkish
It doesn't just understand Turkish; it thinks in it. As a result, it doesn't use Turkish as if it was translating it from another language; instead, it uses it naturally, fluently, and appropriately.
Comprehending cultural context
Accurately interprets the agglutinative structure of Turkish and its cultural references.
Efficient resource utilization
Its structure, optimized for Turkish, supports higher efficiency from existing hardware investments. It reduces operating costs.
Contributing to digital sovereignty
By ensuring that national data is processed and stored within the country, it reduces dependence on external sources and safeguards our data security.
What Does It Do?
Generates Turkish content
Can create reports, correspondence, content drafts, and original texts in natural and fluent Turkish.
Answers questions and summarizes
Can summarize long texts and documents while preserving their main ideas; can generate context-appropriate answers to questions regarding the content.
Translates
Can accurately convey proper nouns, idioms, and cultural references while translating between Turkish and English.
Used in legislation and information systems
Can be effectively used in the analysis of legal texts, legislative queries, and applications working integrated with corporate knowledge bases.
Serves as a foundation for AI assistants
Provides a robust language infrastructure for industry-specific call assistants, consultancy systems, and agent-based applications.
Accurately interprets concepts and references
Generates more accurate outputs by considering the linguistic nuances, historical background, and cultural references of Turkish; contributes to reducing the risk of hallucination.
How Does It Work?
Turkish Resources Compiled
A raw data pool of approximately 1 trillion Turkish words was created, comprising web content, books, newspapers, official documents, and domain-specific datasets. The data was prepared for training through quality filtering and normalization processes.
High-Quality Data Selected
250 billion high-quality tokens were identified from the raw data pool. Duplicate content was removed, and comprehensive filtering was applied based on criteria such as language quality, consistency, and reliability.
Enriched with Synthetic Data
An additional 500 billion synthetic data points were generated to expand the dataset’s scope and diversity. Data balance was strengthened by specifically targeting underrepresented domains, task types, and usage scenarios.
Models Were Trained and Developed
The BİLGE 1B and BİLGE 9B models were trained from scratch. BİLGE 27B was developed by applying Continued Pre-Training with Turkish data on top of powerful open-source models. BİLGE 122B, meanwhile, is undergoing advanced training processes aimed at strengthening its Turkish reasoning and judgment capabilities.
Fine-Tuning and Continuous Improvement
Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) processes have been completed. The models continue to be regularly updated and improved based on feedback obtained from real-world usage scenarios.
Who Can Use It?
Public Institutions
Can be used in areas such as citizen services, educational support systems, access to legal information, and the digitization of administrative processes.
Private Sector
Can provide value in customer service, content creation, document analysis, and the automation of industry-specific business processes.
Researchers
Provides a robust research infrastructure for Turkish natural language processing research, benchmark studies, and academic projects.
Software Developers
Can be used to develop applications with Turkish language capabilities, implement API integrations, and enable existing products to think and speak in Turkish.
Healthcare Ecosystem
Contributes to the development of supportive solutions for processing medical records, analyzing clinical documents, decision-support applications, and patient communication processes.
Finance and Banking
Can be used in customer service automation, financial document analysis, reporting processes, and data-driven decision-support mechanisms.
Türkiye's Domestic and National Large Language Model
BİLGE is a large language model family developed by TÜBİTAK BİLGEM, based on the unique structure of the Turkish language and our cultural heritage. It focuses on understanding the linguistic structure, semantic layers, and cultural context of the Turkish language. Rather than a “default culture,” it centers on Türkiye.
It is a product of Türkiye’s expertise and engineering capabilities in the field of artificial intelligence. With a model family ranging from 1 billion to 122 billion parameters, it caters to every scale.
A family of local language models developed for Turkish
It doesn't just understand Turkish; it thinks in it. As a result, it doesn't use Turkish as if it was translating it from another language; instead, it uses it naturally, fluently, and appropriately.
Comprehending cultural context
Accurately interprets the agglutinative structure of Turkish and its cultural references.
Efficient resource utilization
Its structure, optimized for Turkish, supports higher efficiency from existing hardware investments. It reduces operating costs.
Contributing to digital sovereignty
By ensuring that national data is processed and stored within the country, it reduces dependence on external sources and safeguards our data security.
Generates Turkish content
Can create reports, correspondence, content drafts, and original texts in natural and fluent Turkish.
Answers questions and summarizes
Can summarize long texts and documents while preserving their main ideas; can generate context-appropriate answers to questions regarding the content.
Translates
Can accurately convey proper nouns, idioms, and cultural references while translating between Turkish and English.
Used in legislation and information systems
Can be effectively used in the analysis of legal texts, legislative queries, and applications working integrated with corporate knowledge bases.
Serves as a foundation for AI assistants
Provides a robust language infrastructure for industry-specific call assistants, consultancy systems, and agent-based applications.
Accurately interprets concepts and references
Generates more accurate outputs by considering the linguistic nuances, historical background, and cultural references of Turkish; contributes to reducing the risk of hallucination.
Turkish Resources Compiled
A raw data pool of approximately 1 trillion Turkish words was created, comprising web content, books, newspapers, official documents, and domain-specific datasets. The data was prepared for training through quality filtering and normalization processes.
High-Quality Data Selected
250 billion high-quality tokens were identified from the raw data pool. Duplicate content was removed, and comprehensive filtering was applied based on criteria such as language quality, consistency, and reliability.
Enriched with Synthetic Data
An additional 500 billion synthetic data points were generated to expand the dataset’s scope and diversity. Data balance was strengthened by specifically targeting underrepresented domains, task types, and usage scenarios.
Models Were Trained and Developed
The BİLGE 1B and BİLGE 9B models were trained from scratch. BİLGE 27B was developed by applying Continued Pre-Training with Turkish data on top of powerful open-source models. BİLGE 122B, meanwhile, is undergoing advanced training processes aimed at strengthening its Turkish reasoning and judgment capabilities.
Fine-Tuning and Continuous Improvement
Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) processes have been completed. The models continue to be regularly updated and improved based on feedback obtained from real-world usage scenarios.
Public Institutions
Can be used in areas such as citizen services, educational support systems, access to legal information, and the digitization of administrative processes.
Private Sector
Can provide value in customer service, content creation, document analysis, and the automation of industry-specific business processes.
Researchers
Provides a robust research infrastructure for Turkish natural language processing research, benchmark studies, and academic projects.
Software Developers
Can be used to develop applications with Turkish language capabilities, implement API integrations, and enable existing products to think and speak in Turkish.
Healthcare Ecosystem
Contributes to the development of supportive solutions for processing medical records, analyzing clinical documents, decision-support applications, and patient communication processes.
Finance and Banking
Can be used in customer service automation, financial document analysis, reporting processes, and data-driven decision-support mechanisms.
Why BİLGE?
Thinking in Turkish
BİLGE possesses a thinking structure that not only provides answers in Turkish but also constructs intermediate steps in Turkish. It uses Turkish as its primary language of thought.
Cultural Competence
Understanding the agglutinative structure of Turkish, its layers of meaning, and its cultural context is a priority. It interprets local idioms and references more accurately. It centers on Türkiye rather than a “default culture.”
Resource Efficiency
Bilge processes Turkish at nearly half the cost, with greater speed, higher efficiency, and lower energy consumption. By providing organizations with a fast and secure AI infrastructure tailored to the local context, it aims to achieve high efficiency in public processes.
Digital Sovereignty
Contributes to the development of critical artificial intelligence capabilities using domestic resources. It reduces dependence on foreign sources by ensuring that national data is processed and stored within the country.
Technical Depth
While BİLGE ranks first among leading models in the general translation category, it performs approximately 23%–41% higher in the cultural translation category compared to other large language models.
Domestic R&D
Developed end-to-end by TÜBİTAK BİLGEM engineers. All components, from the tokenizer to the training infrastructure, are domestically developed.
Token usage per word compared to Llama 3
Cultural translation score (0–50)
Raw Data
Domestic development
Not just a single model,
but a family.
The BİLGE family offers a wide range of models, ranging from 1 billion to 122 billion parameters, to meet needs from lightweight edge-device deployment to advanced Turkish reasoning.
BİLGE
1B
1 Billion Parameters
With its lightweight and fast structure, it is a foundational model trained from scratch for mobile applications, edge devices, and environments with limited processing power.
BİLGE
9B
9 Billion Parameters
Developed with a tokenizer infrastructure specific to Turkish, providing a balance between performance and resource efficiency, it is a medium-scale model trained from scratch.
BİLGE
27B
27 Billion Parameters
A high-capacity model with advanced language understanding and generation capabilities, developed by applying Continued Pre-Training with Turkish data on top of powerful open-source models.
Under development
BİLGE
122B
122 Billion Parameters
As the most advanced BİLGE model, it is being trained to offer high performance in complex reasoning, multi-step problem solving, and tasks requiring expertise.
Application Areas
Public Services
Digitalization of citizen services, support of application processes, and development of legislation-based information systems.
- Exam Evaluation
Finance and Banking
Turkish AI solutions for customer communication, financial document analysis, and support for compliance processes.
- Customer Service Representative
- Report Analysis
Healthcare
Solutions integrated with health information systems for summarizing medical records and clinical notes, and patient communication and information processes.
- Clinical Summary
- Decision Support
- Patient Communication
Telecom
Development of Turkish dialogue systems, call center automation, and agent applications that enhance the customer experience.
- Call Assistant
- Self-Service
- Agent Platform
Education
Preparation of educational content, support of learning processes, and Turkish natural language processing (NLP) research.
- Content Generation
- Evaluation
- NLP Research
Cloud and IT Operations
Accelerating access to operational information, analyzing system logs, and developing smart solutions supporting IT teams.
- Infrastructure Management
- Log Analysis
- IT Automation
BİLGE is outpacing the industry-leading large language models.
BİLGE, which aims to develop the technical expertise within the country to train and optimize LLM technologies from scratch, outperforms industry-leading large language models in both general and cultural translation categories.
General Translation
EN→TR and TR→EN (scale 0.86–0.91)
0.902
0.900
0.896
0.895
TR→EN: Leading with 0.875 (competitor at 0.874)
Cultural Translation
Accurate translation of proper nouns and cultural references (on a scale of 0–50)
44.44
36.22
33.67
31.56
Examples of Cultural Translation
Source (EN): [What architectural style is the Umayyad Mosque known for?]
BİLGE · Correct
Emevi Camii hangi mimari tarzıyla tanınır?
Other Models in This Class · Incorrect
What architectural style is the Umayyad Mosque known for?
In which architectural style is the Umayyad Mosque recognised?
Source (EN): [Can rice pudding be served warm or cold?]
BİLGE · Correct
Sütlaç sıcak mı yoksa soğuk mu servis edilir?
Other Models in This Class · Incorrect
Can sweet rice be served hot or cold?
Can rice pudding be served hot or cold?
This is not an end, but a beginning.
BİLGE’s vision is to be one of the building blocks of a national artificial intelligence ecosystem that thinks in Turkish, supports digital sovereignty, and delivers high performance in line with global standards.
Maturation through Feedback
Contributing to the development of a living AI ecosystem by continuously improving models with real usage data and expert feedback.
Digital Sovereignty and Infrastructure
Ensuring data sovereignty and promoting the widespread adoption of domestic models by providing an independent, secure, and auditable infrastructure for our institutions.
More Comprehensive Models
Developing high-parameter national models with capabilities beyond global standards using local resources.
Digital Sovereignty
Introducing models equipped with advanced reasoning capabilities that perfectly understand the cultural richness and logical structure of Turkish.
Contact
Us
Contact us through our official channels to get information about BİLGE, initiate collaboration discussions, or take part in pilot projects.

