Quick experiment I ran because I was curious and had an afternoon free. I took ten sentences from a piece I’d written and put them through ten different free sentence rewrite tools. Then I looked at the outputs.
What I found: six of the ten tools produced outputs I’d call functionally equivalent, meaning same meaning, different words. Two produced outputs where something was lost or altered in meaning, subtle but real. One produced an output that was noticeably better than my original. One produced something I couldn’t use.
The 1-in-10 hit rate on genuinely improving the original was lower than I expected. The 2-in-10 rate of subtle meaning drift was higher. The rest were basically lateral moves, different, not better.
For the use case of ‘rewrite my sentence to improve it,’ free tools are more miss than hit. For ‘rewrite my sentence to make it different,’ they’re more reliable. Those are different use cases and knowing which one you actually need changes whether the tool is worth using.
this is actually a useful data point. i kind of assumed free rewrite tools were either good or useless. ‘mostly lateral with occasional meaning drift’ is a more specific and useful failure mode to know about.
The meaning drift finding is the one I’d flag as most important. A tool that produces a plausible-sounding different sentence but subtly shifts the meaning is more dangerous than a tool that produces an obviously bad sentence, because the bad one you catch and the shifted one you might not.
The ‘different not better’ category is most of what free rewrite tools do and it’s worth being clear-eyed about that. If different is what you need, for variety in content, for freshness in repeated phrasing, these tools work. If better is what you need, they’re basically a coin flip.
The use case distinction you’re drawing is important in professional settings too. Clients often ask for a ‘rewrite’ and mean ‘make this better.’ A tool that produces a lateral move will satisfy neither of you. Getting clear on which kind of rewrite is needed before choosing the tool saves a lot of back-and-forth.
The 1-in-10 that was genuinely better is interesting. I’d want to know what made it better. Was it a structural change, a word choice improvement, a rhythm thing? If you can reverse-engineer what the tool did in that case, you might be able to prompt for that outcome more reliably.