Rendered at 17:26:47 GMT+0000 (Coordinated Universal Time) with Cloudflare Workers.
sligbad 38 minutes ago [-]
Really refreshing read. This feels glaring in so many of these, and the methods to get things to "behave" of just slapping additional markdown prompts at various levels is both silly and ineffective.
Leynos 9 minutes ago [-]
I'd start with something research focused like Undermind or Elicit. Although I don't think that the author is comfortable with using a tool that isn't produced by the model lab.
The planning model for tool use sounds something like CaMeL, which someone should really try implementing in a product.
awakeasleep 45 minutes ago [-]
I have been dwelling on the "No First-Person Output" problem.
I fully agree with the author's point that it's an incoherent interface for a tool. But more than that, it's a constant irritating reminder to me that these LLMs aren't actually thinking or synthesizing new ideas. The LLM is fundamentally not a person, and does not have a human's context, so representing itself with human pronouns and speech patterns is fundamentally contradictory and inaccurate. Author gets into that with the apologies, but once you start noticing it, it's everywhere.
If these programs were actually capable of thinking, and committed to veractiy, they would represent themselves in a new way, and it would be insightful and interesting. We the users wouldn't have comfortable and misleading language masking the 'alien intelligence' and it would be a weird adjustment, but we would be adjusting instead of pretending.
mrweasel 1 hours ago [-]
Whatever it is, it won't be sold as an AI product. Coding and writing tools are probably the easiest to predict. An AI hiding in IntelliSense popping up and warning you that your lacking the proper exception handling, that you're leaking memory and offers to add the missing code, is already doable. Just don't label it as AI, it's realtime security screening for your code.
Or writing an article in Word or Google Docs, having a built in fact-checker akin to the spell/grammar checker is clearly useful. Pink squiggly line, your facts are incorrect, click to fix. Hell built that thing into Facebook or X. Again, it's not completely out of the question to add that right now and have it add the correct sources.
LLMs are clearly useful, but they aren't really a product, they are an engine you can put into other things.
AvAn12 3 minutes ago [-]
Would be lovely, but many issues. 1. If LLMs can't reliably produce factual information today, how can they check if any statement is factual? 2. What about writing that is not recounting facts? e.g. fiction or marketing? 3. who decides what is factual? Does the system give a pass to statements like "full self-driving" or "AGI" or anything accompanied by "the likes of which have never been seen before"?
JKCalhoun 1 hours ago [-]
I don't trust corporations with my data.
So a serious AI product would have to have my data (contexts, conversations) in a "secure enclave". If backed up, it needs to be encrypted.
I want a context and history that, over time, essentially knows everything about me.
It's one of the things that has been rather fascinating about Claude & Co.— he'll come back with things like, "Since you are already familiar with the ESP32…" or, "You already have a heat press from your work with dye sublimation, that will work nicely to set the inks when you screen print t-shirts…"
(Shades of "Diamond Age"… I imagine it helping me recall things when I am in my old age, notice patterns in my life I might want to break free from, etc.)
I don't mind if it's in the cloud if I consent. I always prefer a local-first option. I'm usually attracted to products that offer that. Maybe there are extra features or functionalities if my data is stored on the cloud, but any product that offers a local-first approach is always my preference.
nickdothutton 17 minutes ago [-]
For research tasks I'd like to see labelled branches/traces for the full session/project flow and have the ability to fork from chosen "breakpoints".
k__ 1 hours ago [-]
An AI product should probably start with knowing what model (e.g., arch, version, quant, etc.) you're actually using. Opaque providers make that quite hard.
People are constantly complaining about GPT/Claude constantly changing under their apps without notice.
the__alchemist 1 hours ago [-]
Show me more than 8 items in the recent history list, so I don't have to manually navigate to the same directory repeatedly (Claude)
zzzeek 14 minutes ago [-]
I've definitely seen Claude doing some "double checks" for a lot of its work in more recent versions without my asking it to, and certainly when I use it for important patches, I have another instance of Claude (or sometimes GLM 5.x) do a code review on that patch. Glyph is of course calling for much more prominent UX and gates for these features, good idea.
Razengan 39 minutes ago [-]
It's a shame that YouTubers have dumbed down AI reviews into just "one shooting" random shit that not even they're going to use or play again.
You're not gonna one-shot a real product.
You still have to design the individual elements individually.
Like when trying different models and prompts to generate posters for a hypothetical game, I had to generate a standalone logo first, meticulously and carefully.
You can't just throw them a prompt saying “Make a poster with this and that for a game called MYGAMENAME.”
Even if you have a genie AI you need the darn logo on its own to be able to use it elsewhere.
Similarly you can't just say "Make a fighting game with 900 characters”; you're gonna have to design each individual character on its own.
idle_zealot 27 minutes ago [-]
What you're identifying is a more fundamental bifurcation in why people are interested in AI. Some people have intent, a vision, something they know is possible but lack the technical skills or time to bring into reality. Others want the computer to handle the intent, the technical aspects, the decision-making, the whole process, but be able to go back and specify changes reactively when they don't like something about the result. It seems the latter cohort is much larger.
dist-epoch 49 minutes ago [-]
Would look like a human (robot) you give an access card and point at a desk and it replaces that employee.
empath75 50 minutes ago [-]
Claude Code does most of this stuff now already, in terms of verifications and citations, almost to a fault.
stanfordkid 1 hours ago [-]
The tone of this article is dumb. Of course there are things that can be improved with LLM interfaces, and certainly UX improvements like better citations and grounding can be implemented. But the idea that what has been built "isn't serious" is asinine.
rockskon 46 minutes ago [-]
I'd argue the overwhelming majority of consumer-facing AI products aren't serious products to consumers.
I'm mostly referring to needless AI chatbots shoehorned into various places.
The planning model for tool use sounds something like CaMeL, which someone should really try implementing in a product.
I fully agree with the author's point that it's an incoherent interface for a tool. But more than that, it's a constant irritating reminder to me that these LLMs aren't actually thinking or synthesizing new ideas. The LLM is fundamentally not a person, and does not have a human's context, so representing itself with human pronouns and speech patterns is fundamentally contradictory and inaccurate. Author gets into that with the apologies, but once you start noticing it, it's everywhere.
If these programs were actually capable of thinking, and committed to veractiy, they would represent themselves in a new way, and it would be insightful and interesting. We the users wouldn't have comfortable and misleading language masking the 'alien intelligence' and it would be a weird adjustment, but we would be adjusting instead of pretending.
Or writing an article in Word or Google Docs, having a built in fact-checker akin to the spell/grammar checker is clearly useful. Pink squiggly line, your facts are incorrect, click to fix. Hell built that thing into Facebook or X. Again, it's not completely out of the question to add that right now and have it add the correct sources.
LLMs are clearly useful, but they aren't really a product, they are an engine you can put into other things.
So a serious AI product would have to have my data (contexts, conversations) in a "secure enclave". If backed up, it needs to be encrypted.
I want a context and history that, over time, essentially knows everything about me.
It's one of the things that has been rather fascinating about Claude & Co.— he'll come back with things like, "Since you are already familiar with the ESP32…" or, "You already have a heat press from your work with dye sublimation, that will work nicely to set the inks when you screen print t-shirts…"
(Shades of "Diamond Age"… I imagine it helping me recall things when I am in my old age, notice patterns in my life I might want to break free from, etc.)
People are constantly complaining about GPT/Claude constantly changing under their apps without notice.
You're not gonna one-shot a real product.
You still have to design the individual elements individually.
Like when trying different models and prompts to generate posters for a hypothetical game, I had to generate a standalone logo first, meticulously and carefully.
You can't just throw them a prompt saying “Make a poster with this and that for a game called MYGAMENAME.”
Even if you have a genie AI you need the darn logo on its own to be able to use it elsewhere.
Similarly you can't just say "Make a fighting game with 900 characters”; you're gonna have to design each individual character on its own.
I'm mostly referring to needless AI chatbots shoehorned into various places.