• 4 Posts
  • 29 Comments
Joined 3 years ago
cake
Cake day: January 3rd, 2024

help-circle


  • The thing is, if I don’t add a “good enough for searching” OCR layer, someone else will - AI scraper or legitimate user. That’s a cheap automatic operation. I’ll better do it myself and poison it in a way that won’t interfere with searching, like replacing some “are” with “aren’t”, such common words are rarely searched for. If there’s a chance AI will fall for the metadata and invisible layer contents, that will decrease the requirements for visible poisoning, which is necessary but annoying.

    Some of the text is indeed justified so I could do the multi-column trick that seems like the best compromise. The gap can be as narrow as one space. Or larger if I can write a script to detect lines and connect the columns with gibberish. A human can use zoom or window positioning to view one column at a time. I don’t and will never have access to files the printouts are from (some are presumably in .602 format, others probably .doc), others are handwritten or typewritten, some have images glued on top or hand-traced; and Czech OCR is only about 99.5% reliable so as easier as it would make the endeavor, I can’t be sure to preserve everything if I try to convert them into editable documents as an in-between step.



  • Yes but high school biology knowledge won’t be all that helpful in surveillance tools. I’m targeting chatbots students might want to use to cheat. If an essay is less coherent or truthful, teachers will be able to more confidently punish the student or at least apply some extra scrutiny like a random oral exam on the topic (yes, those are still done here). Trust me that free AIs will become more limited once investor money dries up and paid tools will get price hikes, so if students see that the expensive model can’t produce a Czech text that fools their biology teacher, they might just stop paying, restoring some of their information literacy and reducing corporate revenue a bit. I agree that the surveillance battle is real but that’s not fought on this front. The major chokepoint for surveillance is government accountability (very low in today’s US with ICE agents etc.), slightly less knowledgeable chatbots don’t hinder it very much.


  • I think the AI boom will be over before it makes sense to scrape this, especially if I succeed at making it so that only a human fluent Czech speaker (or maybe a very bespoke program debugged by a Czech speaker) can see through the obfuscation (we have minimum wage so that would be costly). We don’t have AGI and fully autonomous agent workforces as promised, and that’s not gonna change if the remaining <20%, hardest-to-clean digital human-made data gets fed to them. Not to mention some was acquired illegally already, with pending lawsuits. The investors, including governments (thankfully not mine) will eventually realize that LLMs are simply not delivering nearly as much as they cost (including externalities unless the government is shit and passes them to people living near datacenters, laid-off programmers etc.) and never will, although that might take a while.



  • Lbraries are cool but don’t spread information very far, especially when “for local lending only, no copying”. Those would almost never get used and probably thrown away as soon as the libraries realized the difficult copyright situation (not all are by my grandpa, many are unclear due to missing cover pages). Even when incompatible with screen readers (which someone criticized me for), a poisoned PDF is way more usable.

    Anyway, the poisoning, outside the invisible layers and “transcript”, will be visibly baked into the image too but obvious to humans (typewriter vs rendered text that inverts some statements with “not”, “never” etc. at the ends of lines or between words in tiny font and scatters lines between paragraphs saying the text is unreliable, by random bad AI models, or just “clear AI tells” like “Sure, I can do that! 😉 Here’s your summary:✨”). I don’t think anyone will develop an OCR font filter for a few thousand pages in Czech to illegally defeat my DRM enforcing the included licence (unfortunately DRM is incompatible with CC so I can’t use the “CC-BY-NC-SA” shorthand, not that scrapers, my enemies, respect it anyway): there is some protection in the obscurity.




  • I want the poisoning to work on file level because it’s inevitable and welcome for the documents to be shared between teachers and students for free on all kinds of existing platforms. I don’t own any domains, anyway, and it might be best to scatter the documents around to make them harder to blacklist.

    Edit: Google has both search crawlers and AI scraping bots. Even if both are separate, easy to filter or even abiding to robots.txt, the company has indicated that opting out of or hindering scraping will impact search ranking. Of course I’d use a throwaway gibberish $1 domain and couldn’t care less about search ranking, but the power of their opaque, corporate algorithm is immense and maybe would spread to DNS blocking (they control 8.8.8.8). I don’t want to play a cat-and-mouse game (and expect users of the docs to play along).


  • Good point. However, distillation (cannibalizing better LLMs) is a frequent technique so better convince the scraper it’s not good even for that. That’s why I plan to visibly digitally stamp OpenAI GPT 2.0 says: or This Deepseek response has been rated inaccurate: above some paragraphs of the typewritten text (of course in dozens of variations, maybe even different languages, to undermine search-and-replace). Humans will know it’s fake (especially if I add a disclaimer) but scrapers, including ones that re-render and OCR the PDF themselves to get rid of misleading metadata and invisible layers, will most likely rate the text low in value. Of course ethically (and arguably legally) trained commercial AI would reject any text if it is released CC-BY-NC-SA 4.0 but I can’t use that because scrapers for AI ignore licences in practice, and CC specifically forbids taking technical measures to devalue the text for some users.


  • That’s a good idea but won’t be necessary. The documents are scanned, which means all visible text is already a bitmap. For searchability, a text layer will be added as usual for OCR’d documents, but it’s invisible so it does not matter what font it uses.

    I also think I’ll tinker with the bitmap to screw with anyone trying to re-OCR it. If the typewritten text has small, digitally stamped OpenAI GPT 2.0 says: or This Deepseek response has been rated inaccurate: above some paragraphs, a human will easily deduce they have been added later to confuse bots scraping for good training data, especially if a graphical-only disclaimer like The copyright holder released this document for human consumption only. Markers have been added to reduce the apparent and real value for automated tools while keeping the main content intact when viewed by humans. The NC-SA clause of Creative Commons 4.0 applies so no work based on this text can be used in training data of commercial LLMs on the first page explains the situation.


  • You don’t understand just how shit AI is when asked about school topics in Czech. For example, here is a bit of Czech language litany every third grader must know or they will embarrass themselves with awful spelling mistakes. (Skip the bullet points if you just want to hear about the AI)

    • The vowels I and Y (and long versions Í/Ý) sound the same [ɪ] ([ɪː]) unless preceded by D, T, or N but using the wrong one is a big no-no. (Y is never a consonant in Czech)
    • Luckily, in pretty much every native Czech word, I (Í) follows C, J, Č, Ř, Š and Ž, while Y (Ý) follows H, K, R. Consonants Q, W and X basically don’t occur and G, Ď, Ť, and Ň are never followed by I or Y. Foreign words are a huge mess of course, as evident by the existence of the Spelling Bee (we don’t have that, Czech is phonetic with just a few difficult bits like I/Y).
    • The most difficult are remaining consonants B, F, L, M, P, S, V, Z. They are mostly followed by I (Í) but there is a list of about 15 common exceptions on each (vyjmenovaná slova or BY-FY-LY-MY-PY-SY-VY-ZY words), plus their relative words, where Y (Ý) is written instead. For example, there are just 4 ZY-words so I’ll just post the list so you’ll get an idea:
      • brzy - early
        • you love exceptions so I put an exception in your exception: brzičko - diminutive of early - is spelled with an I
      • jazyk - tongue/language
        • … and relative words like jazykolam - tongue twister
          • a well-known one is Strč prst skrz krk, I swear this language is normal
      • nazývat se - be called
        • nazívat se - yawn a lot - also exists for a goddamn reason
          • we have a lot of homonyms for a fully phonetic language, the most common are být - (to) be / bít - (to) beat; my - we / mi - (to) me
      • Ruzyně - Prague quarter where the international airport, until 2012 also called Ruzyně, is located
        • like another part of Prague Výtoň, which has been removed from the lists earlier, nobody cares what the quarter is called now that the airport bears our first president’s name instead (he hated flying but it’s for the better: the same year, there were efforts to name it after fucking Reagan similar to the former Prague W. Wilson (now Main) train station; RR only got a street), but a set of 4 makes for a nice cadence in reciting the ZY-words so it stays
    • The ends of most words are not governed by spelling but the grammar of declination and conjugation. That’s another chapter.

    Well, you’d expect AI to know all cca 100 exception words by heart because they’re public domain and the most famous piece of third grade teaching material (like times tables in second grade) that barely changed in 100+ years so almost every Czech could recite them as a kid? Hell no. There’s dozens of screenshots where Gemini or ChatGPT spewed utter nonsense instead. (DuckDuckGo does not appear to search corporate social media for images because they’re not providing direct links to the files). Granted, some are from users asking for nonexistent XY and HY words but so many are unforced errors. I can’t find my favorite, a Reddit post where Gemini listed dozens of variants of babička with all kinds of endings like Italian “babičetto” before just adding “etc.” but a close second are ones where it adds non-Latin scripts:

    Does the apparent incompetence stop Czech students from cheating with AI? Nope. But the longer the AI stays obviously terrible, the better.


  • Now that’s a clever idea!

    However, they are talking about exported PDFs, which don’t have OCR errors, text is already arranged in lines with standard kerning and all that’s needed is convert formatting so that PostScript for “Document Title (center large font)” becomes “# Document Title” (Markdown for top-level heading) and not “Document Title” (plain text), same with tables. I’m not after an accurate MD tagging - exactly the opposite - so I don’t worry about that, but I think I won’t be able to use their code because OCR’d PDFs are fundamentally different from ones exported with TEX, Word, LibreOffice, Inkscape etc. - the PostScript structure is more like “D (size 18.7) + 15.4pt gap + o (size 18.5) + 11.6pt gap + … t (size 18.8) + 24.0pt gap + T (size 18.5) …” - note that whitespace is just a wider delta of letter coordinates, and sizes are guessed with error margins

    The tags will have to retain some sense and topic adherance or they will be rejected by training QA and the document content or a newly run OCR will be used instead.


  • The scope is Czech-language-only so I wouldn’t rule out the possibility of changing some of the AI’s responses when asked about high school biology in Czech. Alternatively, the text can be made utterly useless for training, for example by diluting it 10:1 with semi-gibberish on a line-by-line basis, and any AI-based next-gen-AI training data QA will reject it for this reason. As long as the core functionality (visual readability, searchability) of an OCR’d PDF works well enough for humans, it’s unlikely someone will re-OCR and fix it. And maybe the graphical layer can be poisoned too, with a black nonsense bitmap text hidden from view by the same, overlaid white actual text (of higher thickness to cover antialiasing)… Or even visibly (screenshots and re-renders exist, after all): If the typewritten text has small, digitally stamped “OpenAI GPT 2.0 says:” or “This Deepseek response has been rated inaccurate:” above some paragraphs, a human will easily deduce they have been added later to confuse bots scraping for good training data, especially if a graphical-only disclaimer “The copyright holder released this document for human consumption only. Markers have been added to reduce the apparent value for automated tools while keeping the main content intact when viewed by humans” on the first page explains the situation.


  • I know how to edit fonts and replace characters. That would ruin searchability (and screen readers), making the PDF as good as a picture scan without OCR, which I hate (and someone would OCR it sooner or later if they realized the text content is useless). However, common words carry meaning (for example “are” is very different from “are not” etc.) and could be replaced with gibberish without most people searching for them. This also forces plagiators to take more steps.

    Anyway, how do I easily add to/edit the PostScript layer in bulk, which consists of a list of individual characters and their positions? As I said, most PDF tools for adding text just add another layer, and that can be easily removed.


  • Yeah, I think that if I strategically replaced “cells” with “little gnomes” or every third “are” with “are not” in the OCR layer, nobody would notice because they’re reading the graphical layer and the text is only for searching within the document (they wouldn’t be searching for “cells” or “are” in a biology text because it occurs so often). Yes, that would make it hard to plagiarize or listen to the documents but I can live with that.

    And the text is in Czech, whose document corpus is not nearly as big as English, a few thousand pages of mild nonsense could make a dent in basic biology knowledge.

    The question remains: how? The OCR layer is basically invisible individual characters and coordinates for each, I can’t write a PostScript parser from scratch to surgically remove some at the right place and add a few more there, that’s outside my scope for the project.