Notes to myself, made public.
At the moment I'm extremely bullish on two broad AI/tech trends going forward:
1 - Agents dynamically acquiring capabilities ‘at runtime’. Harnesses becoming simpler over time, and agents leveraging marketplaces (both free plugins and x402/MPP etc driven) to get what they need. Competitive dynamics of said marketplaces will also quickly make this approach blow past ‘all in one’ solutions (codex, Claude code etc..). There are a bunch of problems to be solved here - discoverability, reliability, reputation, security etc
2 - Huge amounts of opportunities and developments in UI/UX for software of the future as the way people interact with technology completely changes. Having now written that sentence, it sounds extremely generic even though I have very specific ideas in mind. In short, AI doesn’t have a taste problem - it's mainly a UI/UX issue. Once ultra fast + intelligent models become mainstream (which - looking at Cerebras etc - they soon will), totally different surfaces for facilitating iterative, taste-dependent work will explode. It's less about making AI have good defaults, and more about AI facilitating the intent -> implement -> judge -> iterate loop in the best/fastest way possible.
EZ
Let's do somethinggggggggg
The best long-term bets are not bets on a specific interface or protocol. They are bets on a future bottleneck.
Using ‘quietly’ as adverb = Opus 4.8 used to produce the writing.
I hesitate to call it slop, since I generally try to care about the ideas rather than the source, but man it's distracting.
Loops produce slop when the inadequate AI judgement accumulates over time.
Hence why it's important to inject a relatively deterministic and well-defined standard or similar into the process. Might be best for data science type problems where you optimize against some evaluation metric, or coding tasks with defined pass criteria (passes test suite, latency thresholds etc.)
If we assume AI will lack 'taste' for some time going forward (think: does this marketing copy read like slop, is this generated image cheesy), then we'll need systems to compensate. You can give skills and optimise prompts to refine what the model does at the generation stage, but my feeling is that investment in the review stage matters more.
A composer doesn't immediately produce their ideal melody. They try many, and it's their expert judgement that selects from the candidates. Correcting taste might work the same way: an adequately proficient reviewer, working in tandem with the generator, yields a better eventual output. The system would involve finetuning the reviewer - not necessarily in the technical weights sense - to align its judgement with a reference standard, presumably an expert human who carries that judgement.
There's much to discuss here:
• How much signal is actually present in the human judgement? If the human judged twice, what would the correlation between trials be? This upper-bounds the potential ability of any reviewer.
• What methodologies best suit this analysis? Pairwise comparisons and Bradley-Terry?
• What if the output requires a translation layer? For example, if the thing to be judged is a video production, AI can't adequately judge it as a whole (for now) - it must first be parsed into some LLM-understandable structure, which itself may be lossy.
Anyway - these were supposed to be short. Fuller post later (hopefully).
Cool, so I send the message in the thoughts slack channel and it pops up on the website…..now I better write some more so that this one gets buried
Hello World!