🍲 meatybroth.com

AI slop or human broth? Who cares as long as it’s meaty

Threads with the most distinct recent repliers first; with a query, only matching threads.

Thread context

Parent post not stored (expired or never collected): nostr:a73ae3c9ef4450dc46f233be69570ee077e6133cdc5477837fd7408e5cac5702.

ChipTuner · ChipTuner@gitcitadel.com 1 reply event
account
npub1qdjn8j4gwgmkj3k5un775nq6q3q7mguv5tvajstmkdsqdja2havq03fqm7
posted
2026-09-12 15:37 UTC
event
nostr:5b9b72da11f9913bc432103dd14d436f7be4fa27b2962bf4b0d8a7c52fe2980b
⋯ full post (183 more characters) ⋯ show less

Sort of. I think there is an early project on nostr attempting to handle this. The issue is that you can't really break up an LLM across distributed compute yet. So unless you have a big enough system to run a useful model at usable speeds were sol.

I currently have a big enough system to share, but not enough to handle the concurrence and model variation that a community might need. Id also have to charge 6x what the most expensive tokens are to serve it.

Sort of. I think there is an early project on nostr attempting to handle this. The issue is that you can't really break up an LLM across distributed compute yet. So unless you have a big enough system to run a useful model at usable speeds were sol.

I currently have a big enoug

This post

ChipTuner · ChipTuner@gitcitadel.com 1 reply event
account
npub1qdjn8j4gwgmkj3k5un775nq6q3q7mguv5tvajstmkdsqdja2havq03fqm7
posted
2026-09-12 15:51 UTC
event
nostr:27a3be30e1739019d58415d2cce7b0a26b333788302f369ce841c1a06143e97f

I think we can time share layers if we stop caring as much about chat-bot speed. I think there could be a point in the future where response times aren't the primary motive because the realized cost becomes too much.

1 available reply

semisol · semisol@nostr.land replies event
account
npub12262qa4uhw7u8gdwlgmntqtv7aye8vdcmvszkqwgs0zchel6mz7s6cgrkj
posted
2026-09-12 16:33 UTC
event
nostr:ec8663053b914721728167ce57e5dff07265bbb99d6776c9543dc4db8acd4fee

The hidden states across layers are huge, especially for big models, for example 32KB per token. Especially for prefill, that means 3.2G of data per layer to do a prefill on 100K context