Ask HN: (Why) Was the LLM breakthrough useful for images, audio, etc.?

Posted by rogerrogerr 1 day ago

Counter4Comment4OpenOriginal

This has been bothering me for a while. I feel like I have a decent conceptual grasp of what LLMs are doing with written text. But it seems like they also unlocked a bunch of progress in understanding and generating images, audio, and video. I can’t twist my brain into understanding the connection.

Is the boom in generated non-text content also built on LLMs, or is it just correlated with it because a bunch of excitement drove investment into the industry? I’m hoping for an ELI-non-ai-but-cs-major, this has been bothering me for a while.

Comments

Comment by jerlendds 1 day ago

VLM != LLM. Vision language models basically treat text tokens and image tokens the same. Post-training an LLM on images+text can improve its capabilities. Id recommend searching around the keyword VLM to find more resources on how multi-modal AI works.

- https://huggingface.co/blog/vlms

- https://en.wikipedia.org/wiki/Multimodal_learning

Comment by verdverm 1 day ago

transformers were first an image understanding technique, the text processing came later, it's all about the training data and gradient descent, and allegedly attention

Comment by laruss5 23 hours ago

It's actually the other way round - the Transformer architecture was introduced for text (machine translation) in "Attention Is All You Need" (2017). Vision Transformers, which apply it to images, came three years later in 2020: https://arxiv.org/abs/2010.11929

Comment by verdverm 16 hours ago

right, it was not text generation per-se (completion/contemporary understanding) that came first, vision was before that, translation before that

Comment by turtleyacht 1 day ago

[dead]