a half-dozen monkeys with typewriters would eventually write all the books in the British Museum.
Imagine that a billion monkeys with keyboards can do (for this or that corporation, which buys all the available SSD storage supply).
This, by the way, in itself if a bullshit abstract social construction, since the branching factor of a resulting dag (of potential sequences each monkey can produce) is probably larger than “number of atoms in the observable universe” (memes, memes everywhere). You also have to do git merge, which would be an intractable problem.
So, following the Terence’s hysterical post (why, yes, it hurts, when one realizes that a mere mechanical symbol manipulation, which is in itself a good thing, without understanding, could produce “better results” than (You)), it is time do “publish” some premature results, as a good monkey, before corporations collect all the written texts and turn them into DAGs of tokens.
Here is some connection which I has for a long time as in intuitive knowledge, without being able to articulate clearly.
There what underlays any non-bullshit modeling, and unifies the most fundamental structural and algebraic notions which emerge in math and programming .
To accurately model any phenomena, subject to the law of Causality, one has to list or write down all the relevant factors, without missing a single one or adding an imaginary (irrelevant) one.
The “list” can then be “sorted” by a “weight”, which connects to the notion of a partial derivative – how much it “contributes” to the overall process.
This becomes a “weighted sum, which can be thought off as a ;linear combination, and a [column] vector.
When we combine these we get matrices.
When we multiply matrices we compose these linear combinations.
This is how the implicit DAG emerges.
What underlays back-propagation and its vectorized implementation is, thus, discovered, not invented.
The [frequency] probabilistic interpretation “emerges” when we impose a particular set of restrictions on values of the weights.
A “weighted” DAG (updated with each new experience and/or “recall”) is how the brain represents “everything” with synaptic gaps acting as “weights” as well as myelination of the axons, which is also additional “weights”.
Unlike the systematic partial-derivative-based back-propagation, the brain updates its DAG stochastically (which explains the individual differences) but the process converges on all the basic physical constraint of the “shared environment” in which it happen to evolve.
For the LLMs it does not matter, in principle, in which particular “architecture” the [randomly initialized] matrices are composed (multiplied) – it will eventually converge to match the DAG it is being conditioned with.
The notion and the shape (algebraic structure) of a “weighted sum” is what connects everything together. The algorithm is to update the weights in the “right direction”, where the environment is a set of constraints which “define the direction”.
The brain is not a “probabilistic inference machine” (yet another stupid meme), it just updates its structure/representation and it eventually “converges”.
The notion that the “seeking” behavior is driven by the “gap” between the desired state and the “current” state is a universal one, but this is a different (higher) level of the hierarchy biological “abstractions” – several sub-system at work simultaneously. Both the “desired” and the “current” states are just “weights” in different parts of the brain (confirmed by the “re-wiring” experiments) and the updating mechanism is one and the same (this justfollows).
Notice that the Causality Itself, if we strip all the abstract socially constructed bullshit, is an unfolding DAG, in causality, not in “time”, Time is just an abstract notion of an observer.so all the theories taking it as “real” or “fundamental” are, in principle, wrong.
A timescale can be superimposed on a part of the DAG just as a number line (another scale) can be superimposed to from a coordinate system. The key principle is that time as non-existenet as a coordinate axis.