humanizing longer pieces seems to be where every tool falls apart. works fine on 300 words, starts showing problems around 1,000, and by 2,000 words the output has a completely different rhythm from the beginning to the end. is this a known problem and is there a workaround?
Most publicly documented approaches use some combination of perplexity and burstiness. Perplexity measures how predictable the text is relative to a language model’s expectations. Burstiness measures whether sentence complexity varies naturally or stays artificially uniform. AI text tends to be low-perplexity and low-burstiness compared to human writing. Disagreement between tools usually reflects different baseline models and training data, not just noise, though noise contributes too.