Skip to content
Sahn-ı BA; The Courtyard of Basmalah
TREN
SAHN.BA

BA Internet · From data to models

A dataset does not carry the whole world; a model is not truth itself.

Data is selected, named, left incomplete and shaped by context. A model learns patterns from this human-made representation and produces new outputs. BAI does not treat data as ownerless raw material or a model as a mind with authority to judge. It keeps sources, rights, privacy, omissions and human responsibility visible at every stage.

Source · Selection · Permission · Context · Incomplete representation · Correction

A basic distinction

Data is a trace; a dataset is a selected representation; a model produces probabilities and outputs through that representation.

When we reduce a person to a few behavioural records, a city to sensor readings, a language to text found online or soil to measured values, we may lose the nature and relationships that remain. Knowing that data is not the same as reality, and a model is not the same as truth, is not a technical footnote. It is where responsible judgement begins.

What we can measure is a trace of reality, not the whole of it.

Truth and representation

What exists and the representation we record and calculate do not occupy the same place.

01

People, living beings, communities and places

Carry nature and context

  • Carry meaning and relationships that cannot all be measured
  • Change across time, place and circumstance
  • Carry rights, dignity and privacy
  • Have their own voice and right to object
02

Datasets and models

Produce selected representations

  • People choose what is recorded
  • They may be incomplete, old or imbalanced
  • They work toward particular purposes and measures
  • They require fresh review in a new context

Four kinds of record

Read technical terms together with the human choices they carry.

These distinctions make records easier to understand. No record on its own proves reliability or gives permission to release it.

01

Data

A recorded trace of a person, event, text, device or place. Who it belongs to, the context in which it was created and the permission behind it all matter.

02

Dataset

A collection selected, cleaned, labelled and arranged for a purpose. It reflects what was included and what was left out.

03

Algorithm

Steps and rules that determine how an input is processed. People choose which measure to prioritise and which errors to accept.

04

Model

A representation that learns patterns from data and methods to produce predictions, classifications or new outputs. Fluency or success does not give it authority to judge.

From life to a model

Human choices shape every stage. Those choices should remain visible and open to correction.

01Observe and record only what is neededState the real need, the person or community whose data is involved and the purpose from the beginning. Do not collect everything simply because it can be collected.
02Select, name and keep the contextShow what was included or excluded, who applied labels, differences across languages and places and which records remain uncertain.
03Train or build the model for a limited purposeDefine the question the model may answer and the decisions it may not support. Do not quietly move data into a different purpose.
04Test it with real people and in real settingsLook beyond average performance to see who receives errors, which languages or groups are poorly represented and who carries the harm.
05Monitor, correct and stop when necessaryWhen data ages, the setting changes or harm appears, update the model, narrow its authority, delete data or stop using it.

Before using a dataset or model

Answer the same six questions clearly.

  • Where did this record come from?Are the source, date, method of collection, people whose labour made it and history of changes known?
  • Whose rights and permission are involved?Are the rights of people, children, households, communities, creators and institutions—and the terms of use and withdrawal—clear?
  • Who is missing or misrepresented?How do differences in language, place, age, gender, profession, access and other settings remain absent from the data?
  • Where does the model fail?Beyond an accuracy score, are false confidence, fabricated output, discriminatory results and high-risk error examples recorded?
  • Which decision remains human?Is the model’s support clearly separated from the qualified person’s role in checking, accepting, refusing, releasing and carrying responsibility?
  • Is there a path to object and correct?Can people see, correct or delete their data? Can a model error be reported and its use stopped?
When a model is released, the human choices behind it should remain visible. illustration

Open Representation

When a model is released, the human choices behind it should remain visible.

Explain its sources and purpose, what was left out, known errors and uncertainty, decisions for which it must not be used, the person who carries the final judgement, paths for objection and correction, terms for withdrawal and responsibility for care. A technical record then carries not only performance, but also the limits and account of the representation.

A model card is not a shop window; it is a promise of openness to users and people affected.

Sharing and access decisions

Not every dataset or model needs to circulate openly.

SEDD and İHSAN are not approval stamps added at the end. Considering purpose, rights, harm and care together may lead to four different decisions.

01

Do not collect or use it

When there is no real need, or privacy and possible harm cannot be justified, do not create the data or build the model.

02

Keep it private and limited

Set up a narrow use within a household, school, profession or institution, available only to qualified and authorised people.

03

Share it under conditions

Offer controlled access within a defined purpose, age, duration, licence, attribution, human approval and limits on misuse.

04

Open it for shared benefit

When rights, permission, privacy, sources, safety and responsibility for care allow, make it a WAQF candidate through a dated, correctable record.

Shared representation

Turning a community into data without its language, voice and competence is not becoming a MİLLET.

Shared data and model work requires more than a technical team. People represented, native-language contributors, professional practitioners, young people and those responsible for ongoing care should work together.

01

Native language and concepts

A language is more than words to translate. Record concepts, context, history and living use together with people who carry that language.

02

Professional and field competence

From medicine to farming, assess what data reveals and hides together with practitioners who know that field.

03

The voice of people represented

People are not merely data sources. They can see categories built about them, object and take part in correction.

04

Care across generations

Consider who today’s data and model may affect tomorrow. Do not leave updating, deletion and care without a responsible person.

Data and model review

Would you like to examine a dataset or model idea through sources, rights, representation and human judgement?

Shared data and model review Circles within BAI are still in preparation. You can share the real need, data source, people represented and a benefit or risk you have found. This is not a promise of an active model catalogue, data service or programme.