bilge text logo

Evolving with Our Language, Empowered by Our Data

Türkiye's Large Language Model.

About BİLGE

What Is Bilge?

Türkiye's Domestic and National Large Language Model

BİLGE is a large language model family developed by TÜBİTAK BİLGEM, based on the unique structure of the Turkish language and our cultural heritage. It focuses on understanding the linguistic structure, semantic layers, and cultural context of the Turkish language. Rather than a “default culture,” it centers on Türkiye.

It is a product of Türkiye’s expertise and engineering capabilities in the field of artificial intelligence. With a model family ranging from 1 billion to 122 billion parameters, it caters to every scale.

A family of local language models developed for Turkish

It doesn't just understand Turkish; it thinks in it. As a result, it doesn't use Turkish as if it was translating it from another language; instead, it uses it naturally, fluently, and appropriately.

Comprehending cultural context

Accurately interprets the agglutinative structure of Turkish and its cultural references.

Efficient resource utilization

Its structure, optimized for Turkish, supports higher efficiency from existing hardware investments. It reduces operating costs.

Contributing to digital sovereignty

By ensuring that national data is processed and stored within the country, it reduces dependence on external sources and safeguards our data security.

Generates Turkish content

Can create reports, correspondence, content drafts, and original texts in natural and fluent Turkish.

Answers questions and summarizes

Can summarize long texts and documents while preserving their main ideas; can generate context-appropriate answers to questions regarding the content.

Translates

Can accurately convey proper nouns, idioms, and cultural references while translating between Turkish and English.

Used in legislation and information systems

Can be effectively used in the analysis of legal texts, legislative queries, and applications working integrated with corporate knowledge bases.

Serves as a foundation for AI assistants

Provides a robust language infrastructure for industry-specific call assistants, consultancy systems, and agent-based applications.

Accurately interprets concepts and references

Generates more accurate outputs by considering the linguistic nuances, historical background, and cultural references of Turkish; contributes to reducing the risk of hallucination.

Turkish Resources Compiled

A raw data pool of approximately 1 trillion Turkish words was created, comprising web content, books, newspapers, official documents, and domain-specific datasets. The data was prepared for training through quality filtering and normalization processes.

High-Quality Data Selected

250 billion high-quality tokens were identified from the raw data pool. Duplicate content was removed, and comprehensive filtering was applied based on criteria such as language quality, consistency, and reliability.

Enriched with Synthetic Data

An additional 500 billion synthetic data points were generated to expand the dataset’s scope and diversity. Data balance was strengthened by specifically targeting underrepresented domains, task types, and usage scenarios.

Models Were Trained and Developed

The BİLGE 1B and BİLGE 9B models were trained from scratch. BİLGE 27B was developed by applying Continued Pre-Training with Turkish data on top of powerful open-source models. BİLGE 122B, meanwhile, is undergoing advanced training processes aimed at strengthening its Turkish reasoning and judgment capabilities.

Fine-Tuning and Continuous Improvement

Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) processes have been completed. The models continue to be regularly updated and improved based on feedback obtained from real-world usage scenarios.

Public Institutions

Can be used in areas such as citizen services, educational support systems, access to legal information, and the digitization of administrative processes.

Private Sector

Can provide value in customer service, content creation, document analysis, and the automation of industry-specific business processes.

Researchers

Provides a robust research infrastructure for Turkish natural language processing research, benchmark studies, and academic projects.

Software Developers

Can be used to develop applications with Turkish language capabilities, implement API integrations, and enable existing products to think and speak in Turkish.

Healthcare Ecosystem

Contributes to the development of supportive solutions for processing medical records, analyzing clinical documents, decision-support applications, and patient communication processes.

Finance and Banking

Can be used in customer service automation, financial document analysis, reporting processes, and data-driven decision-support mechanisms.

Türkiye's Domestic and National Large Language Model

BİLGE is a large language model family developed by TÜBİTAK BİLGEM, based on the unique structure of the Turkish language and our cultural heritage. It focuses on understanding the linguistic structure, semantic layers, and cultural context of the Turkish language. Rather than a “default culture,” it centers on Türkiye.

It is a product of Türkiye’s expertise and engineering capabilities in the field of artificial intelligence. With a model family ranging from 1 billion to 122 billion parameters, it caters to every scale.

A family of local language models developed for Turkish

It doesn't just understand Turkish; it thinks in it. As a result, it doesn't use Turkish as if it was translating it from another language; instead, it uses it naturally, fluently, and appropriately.

Comprehending cultural context

Accurately interprets the agglutinative structure of Turkish and its cultural references.

Efficient resource utilization

Its structure, optimized for Turkish, supports higher efficiency from existing hardware investments. It reduces operating costs.

Contributing to digital sovereignty

By ensuring that national data is processed and stored within the country, it reduces dependence on external sources and safeguards our data security.

Generates Turkish content

Can create reports, correspondence, content drafts, and original texts in natural and fluent Turkish.

Answers questions and summarizes

Can summarize long texts and documents while preserving their main ideas; can generate context-appropriate answers to questions regarding the content.

Translates

Can accurately convey proper nouns, idioms, and cultural references while translating between Turkish and English.

Used in legislation and information systems

Can be effectively used in the analysis of legal texts, legislative queries, and applications working integrated with corporate knowledge bases.

Serves as a foundation for AI assistants

Provides a robust language infrastructure for industry-specific call assistants, consultancy systems, and agent-based applications.

Accurately interprets concepts and references

Generates more accurate outputs by considering the linguistic nuances, historical background, and cultural references of Turkish; contributes to reducing the risk of hallucination.

Turkish Resources Compiled

A raw data pool of approximately 1 trillion Turkish words was created, comprising web content, books, newspapers, official documents, and domain-specific datasets. The data was prepared for training through quality filtering and normalization processes.

High-Quality Data Selected

250 billion high-quality tokens were identified from the raw data pool. Duplicate content was removed, and comprehensive filtering was applied based on criteria such as language quality, consistency, and reliability.

Enriched with Synthetic Data

An additional 500 billion synthetic data points were generated to expand the dataset’s scope and diversity. Data balance was strengthened by specifically targeting underrepresented domains, task types, and usage scenarios.

Models Were Trained and Developed

The BİLGE 1B and BİLGE 9B models were trained from scratch. BİLGE 27B was developed by applying Continued Pre-Training with Turkish data on top of powerful open-source models. BİLGE 122B, meanwhile, is undergoing advanced training processes aimed at strengthening its Turkish reasoning and judgment capabilities.

Fine-Tuning and Continuous Improvement

Supervised Fine-Tuning (SFT) and Reinforcement Learning with Human Feedback (RLHF) processes have been completed. The models continue to be regularly updated and improved based on feedback obtained from real-world usage scenarios.

Public Institutions

Can be used in areas such as citizen services, educational support systems, access to legal information, and the digitization of administrative processes.

Private Sector

Can provide value in customer service, content creation, document analysis, and the automation of industry-specific business processes.

Researchers

Provides a robust research infrastructure for Turkish natural language processing research, benchmark studies, and academic projects.

Software Developers

Can be used to develop applications with Turkish language capabilities, implement API integrations, and enable existing products to think and speak in Turkish.

Healthcare Ecosystem

Contributes to the development of supportive solutions for processing medical records, analyzing clinical documents, decision-support applications, and patient communication processes.

Finance and Banking

Can be used in customer service automation, financial document analysis, reporting processes, and data-driven decision-support mechanisms.

Why BİLGE?

Thinking in Turkish

BİLGE possesses a thinking structure that not only provides answers in Turkish but also constructs intermediate steps in Turkish. It uses Turkish as its primary language of thought.

Cultural Competence

Understanding the agglutinative structure of Turkish, its layers of meaning, and its cultural context is a priority. It interprets local idioms and references more accurately. It centers on Türkiye rather than a “default culture.”

Resource Efficiency

Bilge processes Turkish at nearly half the cost, with greater speed, higher efficiency, and lower energy consumption. By providing organizations with a fast and secure AI infrastructure tailored to the local context, it aims to achieve high efficiency in public processes.

Digital Sovereignty

Contributes to the development of critical artificial intelligence capabilities using domestic resources. It reduces dependence on foreign sources by ensuring that national data is processed and stored within the country.

Technical Depth

While BİLGE ranks first among leading models in the general translation category, it performs approximately 23%–41% higher in the cultural translation category compared to other large language models.

Domestic R&D

Developed end-to-end by TÜBİTAK BİLGEM engineers. All components, from the tokenizer to the training infrastructure, are domestically developed.

3 vs 0

Token usage per word compared to Llama 3

0

Cultural translation score (0–50)

0 Trillion

Raw Data

% 0

Domestic development

Not just a single model,

but a family.

The BİLGE family offers a wide range of models, ranging from 1 billion to 122 billion parameters, to meet needs from lightweight edge-device deployment to advanced Turkish reasoning.

BİLGE

1B

1 Billion Parameters

With its lightweight and fast structure, it is a foundational model trained from scratch for mobile applications, edge devices, and environments with limited processing power.

BİLGE

9B

9 Billion Parameters

Developed with a tokenizer infrastructure specific to Turkish, providing a balance between performance and resource efficiency, it is a medium-scale model trained from scratch.

BİLGE

27B

27 Billion Parameters

A high-capacity model with advanced language understanding and generation capabilities, developed by applying Continued Pre-Training with Turkish data on top of powerful open-source models.

Under development

BİLGE

122B

122 Billion Parameters

As the most advanced BİLGE model, it is being trained to offer high performance in complex reasoning, multi-step problem solving, and tasks requiring expertise.

Application Areas

Public Services

Digitalization of citizen services, support of application processes, and development of legislation-based information systems.

Finance and Banking

Turkish AI solutions for customer communication, financial document analysis, and support for compliance processes.

Healthcare

Solutions integrated with health information systems for summarizing medical records and clinical notes, and patient communication and information processes.

Telecom

Development of Turkish dialogue systems, call center automation, and agent applications that enhance the customer experience.

Education

Preparation of educational content, support of learning processes, and Turkish natural language processing (NLP) research.

Cloud and IT Operations

Accelerating access to operational information, analyzing system logs, and developing smart solutions supporting IT teams.

BİLGE is outpacing the industry-leading large language models.

BİLGE, which aims to develop the technical expertise within the country to train and optimize LLM technologies from scratch, outperforms industry-leading large language models in both general and cultural translation categories.

General Translation

EN→TR and TR→EN (scale 0.86–0.91)

Bilge-sft

0.902

translategemma

0.900

qwen3.5-it

0.896

gemma-3-it

0.895

TR→EN: Leading with 0.875 (competitor at 0.874)

Cultural Translation

Accurate translation of proper nouns and cultural references (on a scale of 0–50)

Bilge-sft

44.44

translategemma

36.22

qwen3.5-it

33.67

gemma-3-it

31.56

Examples of Cultural Translation

Source (EN): [What architectural style is the Umayyad Mosque known for?]

BİLGE · Correct

Emevi Camii hangi mimari tarzıyla tanınır?

Other Models in This Class · Incorrect

What architectural style is the Umayyad Mosque known for?

In which architectural style is the Umayyad Mosque recognised?

Source (EN): [Can rice pudding be served warm or cold?]

BİLGE · Correct

Sütlaç sıcak mı yoksa soğuk mu servis edilir?

Other Models in This Class · Incorrect

Can sweet rice be served hot or cold?

Can rice pudding be served hot or cold?

This is not an end, but a beginning.

BİLGE’s vision is to be one of the building blocks of a national artificial intelligence ecosystem that thinks in Turkish, supports digital sovereignty, and delivers high performance in line with global standards.

Maturation through Feedback

Contributing to the development of a living AI ecosystem by continuously improving models with real usage data and expert feedback.

Digital Sovereignty and Infrastructure

Ensuring data sovereignty and promoting the widespread adoption of domestic models by providing an independent, secure, and auditable infrastructure for our institutions.

More Comprehensive Models

Developing high-parameter national models with capabilities beyond global standards using local resources.

Digital Sovereignty

Introducing models equipped with advanced reasoning capabilities that perfectly understand the cultural richness and logical structure of Turkish.

Contact

Us

Contact us through our official channels to get information about BİLGE, initiate collaboration discussions, or take part in pilot projects.

BİLGE AI
BİLGE AI