this post was submitted on 17 Oct 2024
713 points (99.6% liked)

Science Memes

11081 readers
2759 users here now

Welcome to c/science_memes @ Mander.xyz!

A place for majestic STEMLORD peacocking, as well as memes about the realities of working in a lab.



Rules

  1. Don't throw mud. Behave like an intellectual and remember the human.
  2. Keep it rooted (on topic).
  3. No spam.
  4. Infographics welcome, get schooled.

This is a science community. We use the Dawkins definition of meme.



Research Committee

Other Mander Communities

Science and Research

Biology and Life Sciences

Physical Sciences

Humanities and Social Sciences

Practical and Applied Sciences

Memes

Miscellaneous

founded 2 years ago
MODERATORS
 
you are viewing a single comment's thread
view the rest of the comments
[–] thevoidzero@lemmy.world 1 points 4 weeks ago

Not just semantics. PDFs doesn't even have segmentations like spaces/lines/paragraph. It's just text drawn at locations the text processor/any other softwares inserted into. Many pdf editor softwares just detect the closeness of the characters to group them together.

And one step further is you can convert text to path, which basically won't even have glyph (characters) info and font info, all characters will just be geometric shapes. In that case you can't even copy the text. OCR is your only choice.

PDF is for finalizing something and printing/sharing without the ability to edit.