Technology

59692 readers

3932 users here now

This is a most excellent place for technology news and articles.

Our Rules

Follow the lemmy.world rules.
Only tech related content.
Be excellent to each another!
Mod approved content bots can post up to 10 articles per day.
Threads asking for personal tech support may be deleted.
Politics threads may be removed.
No memes allowed as posts, OK to post as comments.
Only approved bots from the list below, to ask if your bot can be added please contact us.
Check for duplicates before posting, duplicates may be removed

Approved Bots

founded 1 year ago

MODERATORS

177

ChatGPT bombs test on diagnosing kids’ medical cases with 83% error rate | It was bad at recognizing relationships and needs selective training, researchers say. (arstechnica.com)

submitted 11 months ago by L4s@lemmy.world to c/technology@lemmy.world

24 comments fedilink hide all child comments

ChatGPT bombs test on diagnosing kids’ medical cases with 83% error rate | It was bad at recognizing relationships and needs selective training, researchers say.::It was bad at recognizing relationships and needs selective training, researchers say.

you are viewing a single comment's thread
view the rest of the comments

[–] kromem@lemmy.world 16 points 10 months ago* (last edited 10 months ago)

This is a fucking terrible study.

They compare their results to a general diagnostic evaluation of GPT-4 which scored better and discuss it as relating to the fact it's a pediatric focus.

While largely glossing over the fact they are using GPT-3.5 instead.

GPT-3.5 sucks for any critical reasoning tasks, and this is a pretty worthless study not using the SotA or using best practices in prompting to actually reflect what a production grade deployment of a LLM for pediatric diagnostics would be.

And we really need to stop just spamming upvotes for stuff with little actual worth just because it's a negative headline about AI and that's all the jazz these days.