more stuff for lower cost is good; but for AI to actually achieve lower cost we need to be thinking about intelligence as inclusive of energy efficiency - maximizing highest *value* (quality/cost)- the future is potent models not just enormous models
and what makes models more potent? better architectures a bit, but better data mostly... tbh i think the data preprocessing pipelines are part of the training in a way- filtering n compressing the insights in the pile of tokens
and ppl keep talking abt acquiring more data but actually the key is data *quality* think of it this way- you have a pile of students' ungraded test answers- is this more or less useful for passing the test than just answers from students who have previously received As?