AI Assisted Contributions
- Vous devez vous identifier ou créer un compte pour écrire des commentaires
I know AI is completely a closed source thing at this time, but I have been building up a local AI workflow at home that at least doesn't use the cloud services. Is there a particular stance that Trisquel has on AI-assisted contributions? I am not speaking of fully-automated, where I would not be involved.
I know AI is completely a closed source thing at this time
I do not know. Personal uses of adjusted machine learning models are recent. For that reason, it is better to not draw assertive conclusions.
Nevertheless, in my humble opinion, a set of numbers is not software. It is data. And it is essentially what an adjusted model is. Values for billions of parameters that have been tuned by machine learning. If those numbers are under a free software license (such as Qwen's distilled models, under the Apache 2.0 license), they can be freely used, redistributed and modified, by fine tuning.
To actually use the model (the so-called inference), you need software, running locally to not fall into SaaSS. As you know (since you "have been building up a local AI workflow at home"), there is free software for that. Alpaca ( https://jeffser.com/alpaca/ ) for instance. I have downloaded and used Qwen 3.5 (9B) from Alpaca. No proprietary software is required, if you only use your CPU. You need enough RAM (I have 16 GB) and it is a little slow, but it works fine.
All that said, don't get me wrong: I have many worries regarding LLMs and the likes. They cannot be framed as "AI is completely a closed source thing" though. They relate to the jobs they destroy, the wealth they create that may end up concentrated in the hands of a few capitalists (who are often techno-fascists nowadays...), the vasts amount of energy they require, the disregard for the copylefts on the code in the training datasets, etc.
Is there a particular stance that Trisquel has on AI-assisted contributions?
It is question for quidam and Ark74 to answer.
As long as contributors remain entirely responsible for their contributions, whose quality must remain high, I personally think they should be allowed to use the tools of their preferences. If generative models can boost contributions to free software, they may end up having a positive impact on the movement. On the contrary, prohibiting their usage may make free software technically lose ground against proprietary software that profit from those tools. Also, prohibiting generative models may drive new contributors away. As all tools, LLMs must be mastered. As far as I understand, the code they initially write is often buggy, unmaintainable, etc.
The FSF is working in this area and a draft document is public from 2024:
https://www.fsf.org/news/fsf-is-working-on-freedom-in-machine-learning-applications
Qwen 3.5 does not seem to meet the draft criteria.
https://www.fsf.org/news/fsf-is-working-on-freedom-in-machine-learning-applications says:
After several conversations about the responsibility of the FSF in this discussion, serious work to come to a unanimous conclusion started in May of this year. That work has now concluded, and the working group is currently working to draft the exact text that will form the definition of a free machine learning application.
The year in question is 2024. Do you know where "the exact text" is?
No newer public version of the criteria is available. Still, the existing announcement states that the criteria will require the software, raw training data, and associated scripts to grant users the four freedoms, so we can still evaluate models using those criteria. I don't know of any software/model/training data/script combination that satisfies all of the criteria. It may be impossible to satisfy the criteria at present, but some come closer than others. And that's a fine outcome in my book: Answering what freedom looks like in this context has to include the possibility that nothing currently available is good enough, or else the process is rigged. Unlike the OSI's criteria, where they stated they wouldn't develop criteria that at least one model from today didn't meet. That makes the OSI's outcome seem predetermined.
It seems strange to me that training data would need to be free in order to consider a machine learning application free.
Consider the case of spam detection. If I understand correctly, many email providers use machine learning to detect whether something is spam. To do this, they might collect real spam messages and real legitimate messages and then train a classifier based on that data. Of course, spammers will not willingly license their messages.
If a human were to hand-write a program that detects spam using pre-defined rules, they might use knowledge of some of the patterns that would be encoded in a machine learning model. It seems odd to call the application using a machine learning model nonfree just because a machine was used to figure out the rules instead of a human.
I'm not saying the FSF's statement inconsistent, but it just seems strange to me, given the above possibility.
That said, I am interested in using an LLM that has been trained on a fully free dataset that I can get a copy of. There are some models that aim to train solely on public domain data, like Talkie [1], and they briefly describe where they got a lot of their data, but I don't see a way to download or explore the entire dataset. I don't have the processing power or desire to retrain an LLM from scratch, but I think looking at the dataset could be an interesting way to learn about an LLM, and in the case of models which aim to exclude certain kinds of data, one could attempt to verify that this is the case.
"It seems strange to me that training data would need to be free in order to consider a machine learning application free."
The draft criteria explains why this is the case.
"but I don't see a way to download or explore the entire dataset."
If it truly is not available that makes it nonfree per the criteria.
I don't see where in the linked page [1] that it explains why training data is needed. Is there a separate document with an early draft of the exact text that I am missing? Additionally, I don't see where it says that training data must be provided, unless that's what "the model parameters that represent its training" means, but the word "parameters" makes me think the result of the training rather than the data that led to that result. That page also says all the training data mus be free, but just because something is free doesn't mean it is available. But again, maybe I'm missing the document you are referring to.
The freedom to modify seems like the most likely thing to accidentally be left out of a machine learning application release, but my understanding is that a lack of training data does not make it significantly more difficult to modify a machine learning model.
[1] https://www.fsf.org/news/fsf-is-working-on-freedom-in-machine-learning-applications
The FSF explains this in the "Freedom challenges in ML beyond software". The reasoning is that without the training data, you cannot effectively exercise the freedom to study and modify: "The model parameters are not comprehensible as such by humans, so it is not practical to study or adapt an ML application by analyzing or editing model parameters directly. ... So, in practice, studying and adapting an ML application is usually done, for example, by... analyzing training data, and incrementally training or retraining the model from scratch."
For the other, it's in both the opening paragraph "...which will require the software, *as well as the raw training data and associated scripts*, to grant users the four freedoms" and again in the "Close to a conclusion" section: "...we cannot say a ML application is free unless all its *training data and the related scripts for processing it* respect all users, following the four freedoms."
Further, in the "Freedom may not equal justice" section, the FSF explicitly equates not releasing training data with being nonfree: "It may be that some nonfree ML have valid moral reasons for not releasing training data, such as personal medical data. In that case, we would describe the application as a whole as nonfree."
As a whole. But individual components might still be free, such as the software. But withholding the training data results in the conclusion that they say.
As you wrote, https://www.fsf.org/news/fsf-is-working-on-freedom-in-machine-learning-applications pretends that "in practice, studying and adapting an ML application is usually done, for example, by... analyzing training data, and incrementally training or retraining the model from scratch". It is definitely not true for large models. Almost nobody can afford to train a model from scratch. Take DeepSeek. It always boasts (critics say "lies") that it trains such models for a small fraction of the cost other companies (OpenAI, Anthropic, etc.) pay:
Lastly, we emphasize again the economical training costs of DeepSeek-V3, summarized in Table 1, achieved through our optimized co-design of algorithms, frameworks, and hardware. During the pre-training stage, training DeepSeek-V3 on each trillion tokens requires only 180K H800 GPU hours, i.e., 3.7 days on our cluster with 2048 H800 GPUs. Consequently, our pre-training stage is completed in less than two months and costs 2664K GPU hours. Combined with 119K GPU hours for the context length extension and 5K GPU hours for post-training, DeepSeek-V3 costs only 2.788M GPU hours for its full training. Assuming the rental price of the H800 GPU is $2 per GPU hour, our total training costs amount to only $5.576M.
https://github.com/deepseek-ai/DeepSeek-V3/diffs/0?base_sha=592fd5daf8177b205af11651bbb31a1834a8b0e0&head_user=vaerksted&name=main&pull_number=729&sha1=592fd5daf8177b205af1...
"Only $5.576M". Who can afford that? Advocating that every user must be able to exercise their freedoms by training such models from scratch makes no sense. "In practice, studying and adapting an ML application is usually done" by fine-tuning, contrary to what the FSF wrote. Fine-tuning modifies the model weights using *new* data.
If you disagree that the criteria should require that users deserve to have the training data, I encourage you to send your feedback to the FSF explaining why: https://www.fsf.org/about/contact/email. I imagine that they will use public feedback to inform future versions of the criteria.
I am under the impression that the FSF has already realized the practical impossibility to exercise the "freedom to (re)train from scratch". Probably after receiving feedback from people that are more involved with machine learning than I am (as I wrote below, I have never even tried fine tuning). That would explain why "the exact text that will form the definition of a free machine learning application" has never been published, although it was promised two years ago.
Some do. That is enough for a community of users of a program to control it. Even if none of them can read code, they can contract a developer. To retrain from scratch a large machine learning model, the cost is thousands of times higher. Most importantly, it is unclear whether retraining from scratch actually allows to better modify (to better suit the needs of the community) the model than post-training, which is far cheaper.
I can confirm later drafts, even as recent as a couple weeks ago, still include the training data requirement, so the concept hasn't been abandoned; I can only speculate on the delay. Regarding cost: Free software's about rights, not money, so the "it's too expensive" argument falls flat. What rights do the users *deserve*? The FSF seems to be saying that people should be *able* to do it if they want to or choose to. I expect costs will decrease over time anyway. And, they're trying to write a universal rule for machine learning, not "just" for multi-million-dollar LLMs. For SLM, retraining from scratch on a computer from today is completely viable, making the original training data very relevant.
https://www.gnu.org/philosophy/free-hardware-designs.en.html was written a while ago but what it says seems still valid today:
You can't build and run a circuit design or a chip design in your computer. Constructing a big circuit is a lot of painstaking work, and that's once you have the circuit board. Fabricating a chip is not feasible for individuals today; only mass production can make them cheap enough. With today's hardware technology, users can't download and run a modified version of a widely used digital hardware design, as they could run a modified version of a widely used program. Thus, the four freedoms don't give users today collective control over a hardware design as they give users collective control over a program. That's where the reasoning showing that all software must be free fails to apply to today's hardware technology.
I see similarities and differences with the situation of langage models today. If any individual had the complete training data of most of today's large language models, that individual would be unable to train a model from it. This may change if computing power becomes a lot cheaper but currently we are very far from that. Unlike hardware, the issue isn't mass production because you can still duplicate the weights of a model at not cost, but similarly to hadware, getting those weights out of the complete training data - which are somehow part of the design of the model - is not feasible for individuals today.
It may be feasible for individuals to train smaller models but we need to see how effective it is compared with tuning large models. Certainly, users deserve the freedom to control their computing but we need to find practical ways to achieve that. I have read so many times people saying that the FSF isn't consistent with its approach to freedom because it does not reject hardware made with non-free design but I understand this as a practical approach to focus efforts of what can effectively be reached without a physical barrier, and that is control of the software that runs on your hardware, even though you can't change your hardware.
For language models, the physical barrier is not the same like for hardware manufacturing, so I am not saying it should be treated the same. Reading https://www.fsf.org/news/fsf-is-working-on-freedom-in-machine-learning-applications, it seems to me that the FSF is having a completely practical approach, like it had about hardware, by not calling free large language models with training not freely available but not rejecting their usage and trying to define criteria to improve user control over them (like right to share weights, tune the model and share the tuned model). I am anticipating too much there, let's wait for the exact text from the FSF. You called the FSF a lighthouse, I view it more as a compass, helping to make proper decisions.
I would like to complement Avron's excellent points on the similarities and differences between hardware designs and "training data and the related scripts for processing it" (in the context of machine learning models). Let me first quote the beginning of https://www.gnu.org/philosophy/free-hardware-designs.en.html#reject-nonfree (Avron chose a subsequent excerpt of that same section, entitled "Must We Reject Nonfree Digital Hardware?"):
Is a nonfree digital hardware design an injustice? Must we, for our freedom's sake, reject all digital hardware made from nonfree designs, as we must reject nonfree software?
Due to the conceptual parallel between hardware designs and software source code, many hardware hackers are quick to condemn nonfree hardware designs just like nonfree software. I disagree because the circumstances for hardware and software are different.
Present-day chip and board fabrication technology resembles the printing press: it lends itself to mass production in a factory. It is more like copying books in 1950 than like copying software today.
Freedom to copy and change software is an ethical imperative because those activities are feasible for those who use software: the equipment that enables you to use the software (a computer) is also sufficient to copy and change it.
Besides the insufficiency of home equipment to relearn a large model from its training data (I have already made that point), I notice that, while "the conceptual parallel between hardware designs and software source code" is obvious, I would not say that there is an analog conceptual parallel between "training data and the related scripts for processing it" and software source code.
Indeed, in https://www.gnu.org/philosophy/free-sw.html#make-changes (a section of FSF's free software definition), "source code is defined as the preferred form of the program for making changes in" and, as far as I know, hackers who are into machine learning prefer to make changes to a model by post-training it (what is feasible) and not by relearning it from scratch. As the name says, post-training does not start from the training data, but from the learned weights, using new data. As a consequence, the learned weights better qualify as source code than "training data and the related scripts for processing it".
I was curious to see if maybe some older models had a training data release, so I was looking at information about GPT-2 [1], and I then looked at the Wikipedia page for GPT-1 [2], which I noticed had a link to "Open-source artificial intelligence" [3] which apparently does actually require that training data be available. The Wikipedia page makes it sound like the training data has to be free too, but it seems like maybe the OSI disagrees with that. I looked at a source mentioned on that page [4], which mentioned that the OSI would list some models that meet their definition, and the first one I looked at, "CrystalCoder" [5], seems at a glance to have all training data available [6], but some of the data is nonfree (e.g. CommonCrawl). The OSI has a page listing more models [7].
With full access to the training data, someone with the resources to do so could filter out any nonfree data from the training set and then re-train it. Perhaps some of the models that the OSI considers to be "open-source artificial intelligence" may be small enough that anyone could do this.
[1] https://en.wikipedia.org/wiki/GPT-2
[2] https://en.wikipedia.org/wiki/GPT-1
[3] https://en.wikipedia.org/wiki/Open-source_artificial_intelligence
[4] https://www.technologyreview.com/2024/08/22/1097224/we-finally-have-a-definition-for-open-source-ai/
[5] https://github.com/LLM360/crystalcoder-train
[6] https://huggingface.co/datasets/IFM/CrystalCoderDatasets
[7] https://opensource.org/ai
Yes, the OSI's definition at https://opensource.org/ai/open-source-ai-definition also has requirements for training data but note that it's much weaker and none of it has to be free, and in fact none of it even needs to be sharable or be shared, while under the FSF's criteria it does. I also note that even the OSI calls such data the "preferred form to make modifications to machine-learning systems".
> With full access to the training data, someone with the resources to do so could filter out any nonfree data from the training set and then re-train it.
In theory, but I see a significant practical problem: the amount of effort to audit, for example, The Common Pile's 8TB of text seems impractical. One thing that I'd like to see in the FSF's criteria is a "Commitment to Correct Mistakes" for model and data set maintainers much like endorsed distro maintainers do. In order to recommend a particular model or training data set, the maintainers would make a reasonable effort to keep out non-free junk and to remove it if found.
Did you try fine-tuning? If so, what kind of data did you use for that and did you notice a change when asking certain questions between the answer without and with fine-tuning?
Besides, I am not sure how to master these tools, can fine-tuning do that?
If an LLM was trained with non-free software or leaked pieces of information, there is a risk that it writes code that more or less trivially reproduces pieces of non-free software or takes advantage of the leaked pieces of information. If that code is used for a free software project, the project may be sued by a company arguing that their copyright was violated or that it performed reverse engineering violating their usage conditions.
Even if the LLM would be trained with code under the same license, there is a copyright owner that should be credited anyway. Unless LLMs can be adjusted in such a way that they can identify the source and the license of the code they somehow reproduce and provide the associated copyright notice, it looks difficult to eliminate legal risks if using them for free software projects.
Did you try fine-tuning?
I did not.
I am not sure how to master these tools, can fine-tuning do that?
The freedom to modify certainly participates in controlling the work you achieve with the help of an LLM. That said, I was actually thinking of mere uses. I guess the first step into mastering these tools is to be critical of their output, to not blindly accept it.
If an LLM was trained with non-free software or leaked pieces of information, there is a risk that it writes code that more or less trivially reproduces pieces of non-free software or takes advantage of the leaked pieces of information.
That is indeed a risk. Essentially all (free or proprietary) software companies are taking it today. Hopefully, those who trained the model in the first place (disregarding the licenses) will end up in legal troubles. Not the end users of the model.
>" I have downloaded and used Qwen 3.5 (9B) from Alpaca. No proprietary software is required, if you only use your CPU. You need enough RAM (I have 16 GB) and it is a little slow, but it works fine."
Ok, I think we need to see a @Magic Banana Qwen 3.5/Alpaca how-to/walkthrough. I've been wanting to do the same, but would prefer to learn from your successful actions than having to re-invent the wheel, if possible.
It is really easy. I installed Alpaca from Flathub (the first instruction below installs flatpak on your system; the second instruction adds the floss subset of the Flathub repository; maybe all that as already been done on your system):
$ sudo apt install flatpak
$ flatpak remote-add --subset=floss flathub https://flathub.org/repo/flathub.flatpakrepo
$ flatpak install flathub com.jeffser.Alpaca com.jeffser.Alpaca.Plugins.Ollama
(Warning: even asking for the "floss" subset, as above, applications only aiming to launch proprietary software are proposed on Flathub.)
Launch Alpaca. In it, choose "Manage Models" (in the burger menu or using Ctrl+M), click on the "Available" tab and search the model of your choice. You can directly type its name if you know it.
(Warning: some models labeled as "Cloud" are SaSS, others have unacceptable terms of service, others are too large for your system, and I believe none would satisfy the FSF's current stance.)
Once the model downloaded, go back to the main window. There you can allow the model to use tools such as "Web search" and input a prompt to start a discussion.
(Warning: the integrated Web browser may execute proprietary JavaScript, I believe.)
So, as you can see, there are a number of real freedom pitfalls.
Would you call your approach to getting a local ai as good as it gets regarding free software?
Alpaca relies on Ollama, which is distributed under the Expat (often called MIT) license, which is a free software license. Alpaca itself looks entirely GPLv3-licensed. Some of the models it proposed are unacceptable, from a freedom perspective. In particular, using the models tagged "Cloud" is definitely SaSS. I searched more since my last post. Maybe the Olmo models, which are available in Alpaca, are up to FSF's standards (but good luck trying to exercise in practice the freedom to relearn them from scratch): https://allenai.org/olmo
I would be happy to learn about any freedom issue I would not be aware of.
I've been doing my own research and this looks to face the same issue as others regarding non-free training data, though I can see why people might think it's free.
For training data, the OLMo project created a dataset called Dolma, released under ODC-BY, which looks to be a a permissive attribution- only license): https://opendatacommons.org/licenses/by/1-0/
All supporting code and tools are under Apache 2.0.
The nuance: The Dolma dataset compiles data from sources like Common Crawl, Wikipedia, Project Gutenberg, and academic papers. Common Crawl in particular mixes free and non-free works. The Dolma README says "We are releasing this dataset under the terms of ODC-BY. By using this dataset, you're also bound any license agreements and terms of use of the original data sources."
In principle, if someone were to filter out the non-free portions of Dolma, then re-train OLMo2 (or build a new model from scratch) on just the free stuff, the result could be a model that is entirely based on freely licensed training material. If one applied a free license to the model, combined with the already-free llama.cpp software, this would effectively solve the problem though of course, accomplishing this is far from a trivial task.
To put it in perspective, the Allen Institute for AI (AI2) used "Augusta," which is a 160-node supercomputer cluster provided by Google. Each of the 160 nodes is equipped with 8 H100 GPUs from nvidia, meaning the training process utilized a total of 1,280 GPUs and ran for approximately 30 days of continuous computation on that system. This is a practical barrier to exercising the freedom to modify and rebuild these things but still - perhaps a crowdfunding effort, if large enough, could raise a sufficient amount of money.
Although nvidia is not a good example to use. On Trisquel, the only available free driver for NVIDIA cards is Nouveau, which lacks a working implementation of CUDA - the proprietary software that NVIDIA GPUs depend on for this kind of activity. This effectively rules out NVIDIA hardware without significant reverse engineering work.
At first there appeared to be two promising paths:
1. The Common Corpus and Pleias models: A French organization, Pleias, has released the Common Corpus training data, which is entirely composed of permissively licensed and public domain sources. But what's considered public domain in one area might not be in another. This raises potential issues that The Common Corpus might not be free for everyone, everywhere. Unless the material is very old, achieving a worldwide "free of copyright" status is difficult, which is what CC0 aimed to address.
Their Pleias model family and weights are published under the Apache 2.0 license. Even the training toolchain they use - HuggingFace's Nanotron - is free software but it, too, has a problem with a hard dependency on NVIDIA's proprietary CUDA software. To my knowledge we don't have a free CUDA implementation to use.
2. The Common Pile and Comma models Separately, EleutherAI has released training data called The Common Pile, also built from freely licensed text. The Common Pile should not be confused with with an earlier version called The Pile which should be avoided. Their Comma models and weights trained from The Common Pile are available under Apache 2.0, and they were trained using GPT-NeoX, which is also free software.
GPT-NeoX already has support for AMD's ROCm software. Based on this, the best candidate for now looks like a combination of llama.cpp + the GPT-NeoX toolchain for training + The Common Pile data set.
Even though the situation with AMD and Intel GPUs is somewhat better it still ultimately fails too: Both require non-free firmware to be loaded, which means the ability to train an LLM entirely in freedom (on a GPU) doesn't exist at present despite that we can run models that have already been trained without needing that software.
Thank you for this information, these were the kind of models I was looking for earlier but didn't find. I tried using a derivative (a conversion to the "GGUF" format) of EleutherAI's Comma model which works with llama.cpp, and it's interesting to interact with since it's a base model. I don't think llama.cpp is meant to work with base models, or maybe I'm missing an option that makes it work better. Often the model acts as the user asking another question after answering the first one.
Interestingly, when I asked "Who are you?" is started with "I am an AI language model developed by OpenAI.". I wonder what made it decide to say that. If I had enough storage space for the Comma dataset, then I could grep for "OpenAI" in the training data. The dataset seems to be around 521 GB. Actually, that's small enough I really could download it, as I technically have just enough space on my home server.
Base models are raw text predictors so you interact with it differently. Don't treat it like a chatbot, but do use the text completion aspect by re-phrasing so that it's not a question but an incomplete sentence for it to finish. "What is the capital of France?" becomes "The capital of France is". Or use "few-shot prompting" by providing a few examples of a Q&A format before letting it finish.
Even with good prompting, a base model may not stop after finishing. Look into using a stop sequence. In llama.cpp that's the -r/--reverse-prompt flag. Without it, the model may happily keep going.
The other probably comes from the web having so much LLM outputs copied from ChatGPT that the model learned "I am an AI language model developed by..." is very commonly completed with "OpenAI." That's LLM-generated text contaminating the crawl, which is hard to fully filter even from a carefully-licensed training data.
I thought I would need to use a style like what you said, but it seems like the "chat template" might be getting in the way of that usage sometimes. For example, when I said "An SVG of a pelican riding a bicycle: < svg" (I did not have a space between '<' and 's' in my text, but I've added one to avoid breaking the forum text.), the model said:
Sure! Here's a SVG of a pelican riding a bicycle:
<|im_end|>
In this case, the model stopped outputting after that, but in other cases it did loop until I manually stopped it. I can put "<|im_end|>" as the reverse prompt, and that seems to fix the looping as well as hide the "<|im_end|>" strings, but I think the models may seem less capable than it is because of the chat template, as it's playing the role of an "assistant" despite not being specifically trained to.
I did find a GitHub discussion which suggests using a chat template file [1], but I have not tried that yet.
"Both require non-free firmware to be loaded"
So in the case of NVIDIA the problem starts from the drivers, while in AMD and Intel GPUs cases it comes from the firmware, correct? I'm asking because I believe that the topic of firmware is touchy topic when it comes to free software in this forum, and it appears to me that it boils down to whether or not the firmware can be updated or not, as it happened with BIOS and the reason to start a campaign to develop a free BIOS too.
Unless I mistaken, the case of the BIOS it represents a high risk, since it could potentially be use to spy on the user, can it not? In the case of GPU's firmware, I imagine that the only changes with updates are in performance, not privacy, is it not? Important for sure, but not as a high risk in my opinion.
Also, if a privacy vulnerability is discover on a firmware that can not be change, it will be relatively easy single out. But if it is something constantly being updated, there is not end to what it can do. So I do think there is an importance in whether or not the firmware is 'update-able'. Unless I'm miss remembering, RSM explained like this: say we have a toaster, if it can not be update-able is it really necessary to know how it bakes? if we analyze it and it doesn't have any communication functionality, do we really need to know what it does? If for whatever reason a malicious functionality is discover, it is simply single out as such so that no further harm can be done.
Sorry If I miss remember or miss understood RSM message, any clarification is welcome.
In any case, it would be for sure a good thing if firmware could be available as well but I find it less important and feasible as in the case of drivers.
I think, ironically, the Paul Allen Institute might be moving things in a positive direction.
from https://allenai.org/blog/tulu-3-technical
"This lack of transparency creates challenges for reproducibility and hinders progress in understanding how specific fine-tuning strategies impact model performance. With Tülu 3, we are releasing state-of-the-art post-trained models with every step in the pipeline open – training datasets, data curation tools, data decontamination scripts, training code, evaluation suites, etc. We believe this will both close the gap to closed recipes for post training and act as a foundation for the next chapter of open post-training research."
If indeed every step in the pipeline is open wouldn't that overcome much of fsf's concerns? Or am I misreading things.
That deals with "post-training" (what includes "fine-tuning"), not "training or retraining the model from scratch", as written in https://www.fsf.org/news/fsf-is-working-on-freedom-in-machine-learning-applications
I have been using llama.cpp (software behind ollama) which is FOSS, so my question was more about using the LLMs themselves like Qwen and Deepseek, for example. I happen to know that both of them must have been trained with all types of sources free and non-free. I suppose its okay to contribute to free software with non-free tools. After all, we all use non-free system RAM blobs. :)
I just wanted to poke at this talking-point to see how the Trisquel community acts. I'm just glad I was not immediately cancelled out of existence. Ark once told me long ago to "just do it" so I will. Thanks all.
"After all, we all use non-free system RAM blobs. :)"
I find it sad that people think that free software enthusiast are so closed minded and hard to talk with, specifically the people of this forum. There are people that I truly respect like Magic Banana who I find is a very respectful and eloquent.
I mean sure, no community is perfect, but I don't see a reason to take anything personel in the internet. The most we can do is disagree after laying down our understanding and positions. There is no good on pointing out others personal flaws, but our ideas and opinions on them.
About the case of firmware RAM, I shared my opinion about it in my last post https://trisquel.info/es/forum/ai-assisted-contributions#comment-184117 , is RAM firmware update-able? xD just curious.
In any case, as a free software enthusiast, I don't think we are trying to live in a world of fairytales, I think our purpose is betterment not perfection, that's what I believe we should aspire towards to.
I agree with you. I am actually extremely on-minded and pragmatic about free software. I make my efforts point toward and convert to free software tools and output as much as I can; and I try hard.
The range of topics covered in this forum is also quite open.
About AI: would "SLMs" be of any use in your efforts to contribute to Trisquel?
https://trisquel.info/en/forum/ai-assisted-contributions#comment-184055
Disclaimer: I just learned about "SLMs" while reading this thread. :)
non-free system RAM blobs
Sorry, I am not sure what blob you are referring to. Are you referring to RAM training code executed at startup by the CPU? I understand that on some systems, this is released as free software.
Yes. Which system? My understanding is that there were no free options? I do recall the main issue is whether the blob is updatable or not. I was being a bit cheeky with my "we all use RAM blobs" comment.
I have not checked the source by myself, but I understand that systems using gnuboot should not have any non-free blob to boot, and it is possible to build the arm-trusted firmware and u-boot without blobs for some board like rockpi4. Some experience is shared at https://ixypsilon.net/rock-pi-4-blobless-bootloader/. I also did that on rockpro64, which worked for a while but I now have no time to spend on this.
I don't really see how free (as in freedom) training data is feasible in this current era of LLMs that are trained on pretty much the entire web. Maybe if you distill an open-weight model (like DeepSeek), but that would still rely on the model you're distilling being trained on nonfree data, and I'm not sure how this meaningfully increases your freedom.
Do you mean that SLMs are in fact "distilled" LLMs?
As the S in the acronym says, SLMs are small. That does not tell anything more. Anyway, when a same model (such as Qwen 3.5) is followed by different numbers (such as 9B or 27B, where B stands for billions) then my understanding is that the smaller models are usually distilled versions of the largest one, which is often impossible to run on home equipment (the full version of Qwen 3.5 has 397 billions parameters that do not fit into your RAM).
Yes, we are indeed talking about "Small Language Models" vs. "Large Language Models", where size is measured in number of parameters.
I was originally thinking about those SLMs that are not "distilled" versions of LLMs. Do you know any? They may have been trained for specific tasks, which may not require crawling and wheigting the complete contents of reddit and wikipedia or the whole code base of every software editor.
They indeed do not have to be distilled from bigger ones. One could look at Microsoft's TinyStories for example and have a model with millions of parameters rather than billions for CPU training.
https://news.microsoft.com/source/features/ai/the-phi-3-small-language-models-with-big-potential/
Something of this size seems trainable within hours on CPU.
https://www.reddit.com/r/LocalLLaMA/comments/1r8c6th/flashlm_v4_43m_ternary_model_trained_on_cpu_in_2/
It might be interesting to do some experiments to see how big one can go and still be trainable on CPU within some reasonable amount of time. Whatever "reasonable" is, considering that the training process for OLMo2 ran for 30 days and that seems on the short side. Not that we can make anything that big but, but running for that amount of time. If only there were someone with a Talos II to pack in as much CPU punch as possible during such an experiment...
TinyStories has been generated by an LLM. The link you gave clearly states it:
Then [Microsoft researchers] asked a large language model to create a children’s story using one noun, one verb and one adjective from the list – a prompt they repeated millions of times over several days, generating millions of tiny children’s stories. They dubbed the resulting dataset “TinyStories”.
Why would post-training an LLM give us less control (the whole point of the free software movement) over the resulting model than learning from millions of texts generated by an LLM? The latter is actually a form of knowledge distillation, a type of post-training: https://en.wikipedia.org/wiki/Post-training_of_large_language_models#Distillation
Instead of directly learning to match the probability distributions the LLM computes for the next tokens, you learn it through millions of texts generated according to those same distributions. When the number of generated texts tends to infinity, I believe it is the same information the teacher model transfers to the student model. That way of transferring it is less efficient though. It is used when the teacher model's output distribution is unavailable. Typically, if it an LLM on a remote computer you do not control (aka "the cloud"). In the case of "TinyStories", GPT-3.5 and GPT-4 (behind the respective versions of ChatGPT) generated the stories: https://arxiv.org/abs/2305.07759
I feel like we're getting sidetracked by semantics, but I'll address it briefly so we can get back to the actual point.
While machine learning literature sometimes colloquially calls training on model outputs "black-box distillation," classical knowledge distillation specifically transfers the teacher's internal probability distribution (logits/soft labels). When training on synthetic text (like TinyStories), the model is just doing standard next-token prediction on hard labels. From the perspective of the training algorithm and the hardware, the process is mathematically indistinguishable from training on text written by a human.
That aside, my point was never about whether TinyStories used a philosophically "pure" dataset. I brought it up purely to illustrate model architecture and compute feasibility.
We can't build LLM with free software because LLMs require millions of dollars in proprietary GPU clusters. But SLMs show that small one (tens of millions of parameters instead of hundreds of billions) can be trained from scratch on standard hardware.
So, putting dataset provenance aside for a moment: It would be interesting to do an experiment to see how large of a model we can actually train from scratch with CPU training on accessible, owner-controlled hardware like a dual-socket POWER9 Talos II within a practical timeframe. Even if we're restricted to small parameter counts, just how capable can we make an SLM while keeping the entire stack (code, and dataset) within the bounds of software freedom?
Does anyone with a Talos II want to benchmark something like this? I could use my D16s but, given the same-sized processing window (say, 30 days of training) a Talos II could process more training data and could and let us make an even bigger model than my D16s could ever hope to.
From the perspective of the training algorithm and the hardware, the process is mathematically indistinguishable from training on text written by a human.
The process does not start with the synthetic data, since they are synthesized. It starts with the LLM. In the article you pointed, the section "The role of high-quality data" is all about the importance of synthesizing the learning data with an LLM. In particular, the large excerpt between those two quotes insists that using raw data would not result in good small models:
“Instead of training on just raw web data, why don’t you look for data which is of extremely high quality?” asked Sebastien Bubeck, Microsoft vice president of generative AI research who has led the company’s efforts to develop more capable small language models.
(...)
“The power of the current generation of large language models is really an enabler that we didn’t have before in terms of synthetic data generation,” said Ece Kamar, a Microsoft vice president who leads the Microsoft Research AI Frontiers Lab, where the new training approach was developed.
The function (in the mathematical sense) going from the teacher model (an LLM) to the student model (which can be small) is the whole "process". I believe it is asymptotically indistinguishable from white-box distillation. If the number of texts does not tend to infinity, the resulting student model is worse at imitating the teacher model.
If possible (unlike with OpenAI's models, used to synthesized TinyStories), white-box distillation is preferred to black-box distillation. Not only because it is a more effective way of transferring the knowledge from a teacher model to a student model, but also because it is more efficient.
So, my point is: if black-box distillation is ethically acceptable, then white-box distillation should be ethically acceptable and we had better post-process an existing LLM (whose license lets us do so) than train a model from synthetic or (even more difficult) raw data. Going the harder and worse route you advocate does not give us any more freedom, i.e. any more control over the resulting model. If it actually does (i.e., if I am mistaken), then I would really like to be presented concrete examples.
That aside, my point was never about whether TinyStories used a philosophically "pure" dataset.
So, you consider that a model learned from TinyStories cannot possibly be ethically acceptable, right?
But SLMs show that small one (tens of millions of parameters instead of hundreds of billions) can be trained from scratch on standard hardware.
Not "from scratch", but from an LLM, if you want a reasonably-good SLM, as in the examples you gave. I repeat: the point of Microsoft's article is that using an LLM (not raw data) allows to effectively train an SLM.
My main question remains unanswered: why would post-training an LLM give us less control (the objective of the free software movement) over the resulting model than learning it from scratch?
Clearly, the answer is not training on the *entire* internet. :)
The post seems to have drifted a bit.
> Is there a particular stance that Trisquel has on AI-assisted contributions?
I can't answer the question directly, but regardless of the direction Trisquel chooses, using LLMs don't make other rules go away.
So for instance, if you generate code or data with an LLM, LLMs can produce code or data that come from other places.
According to Software conservancy, at the time when I asked (last summer), the models were trained on mostly free software code or data, but nonfree code or data could still be in the training data (for instance leaked source code that wasn't removed yet from github). This may have changed and/or depend on the models, software used, etc.
But even if an LLM is trained only only free software code/data, at least 2 legal issues remain:
* You could end up combining code that has incompatible free software licenses.
* Most free licenses require you still need to attribute the author and here the LLMs don't tell from where they copied or adapted the code.
So you need to make sure that the generated code/data doesn't meet the Threshold of originality[1], which requires legal skills for both you as the contributor and/or for the people that do review and accept your patches. That kind of legal skills was not something free software had to deal with before, as before the fact that we knew the provenance of the code simplified things a lot.
In addition Software Freedom Conservancy also has guidelines[2] for the project policies and/or how contributors are expected to behave, so for instance whatever you do, you are expected to disclose that your contribution uses LLM as per these guidelines..
At the end of the day Trisquel could also decide to ban LLMs, or to wait for projects like GNU to come up with their policy, which will be made with the help of lawyers, to be on the safe side.
In addition, as I understand, producing code of good quality with LLMs also require a huge amount of skills in project maintenance, and if contributors don't really understand what the LLM is producing that puts more burden on the maintainer, and doesn't necessarily help the contributor acquiring these skills.
So the question here is also if it's really worth the effort as the Trisquel maintainer may also not have the time to look into all that mess to have an opinion on how to deal with LLMs.
[1]https://en.wikipedia.org/wiki/Threshold_of_originality
[2]https://sfconservancy.org/llm-gen-ai/llm-backed-generative-ai-recommendations.html

