
Security News
arXiv Is Rate Limiting Authors Following a Flood of AI Slop Submissions
arXiv now limits authors to two submissions a month as AI slop overwhelms moderators, delays good papers, and sparks debate over applying the limit to everyone.
@codragraph/compress
Advanced tools
Lossless semantic compression for LLM contexts — strip grammar, keep facts, save tokens.
Strip grammar. Keep facts. Save tokens.
- "In order to optimize the database query performance, we should consider implementing an index on the frequently accessed columns..." (70 tokens)
+ "Need fast queries. Check which columns used most. Add index to those columns..." (50 tokens)
= 29% reduction

LLMs excel at filling linguistic gaps. They predict missing grammar, connectives, and structure.
Key insight: We remove only what LLMs can reliably reconstruct.
What we remove (predictable):
What we keep (unpredictable):
Compressed: "Company medium-large. Location Stockholm."
Decompressed: "at a medium-large company based in Stockholm"
↑ grammar added, facts unchanged ↑
LLM-based (best compression, requires OpenAI API):
python3 -m venv venv
source venv/bin/activate
pip install -r requirements.txt
# Set up API key
cp .env.example .env
# Edit .env and add your OpenAI API key
NLP-based (free, offline, multilingual):
python3 -m venv venv
source venv/bin/activate
pip install -r requirements-nlp.txt
python -m spacy download en_core_web_sm # or other language models
MLM-based (free, offline, predictability-aware):
python3 -m venv venv
source venv/bin/activate
pip install -r requirements-mlm.txt
python -m spacy download en_core_web_sm
NLP-based compression (most stable, 15-30% reduction, free, offline):
python nlp.py compress "Your verbose text here"
python nlp.py compress -f input.txt -o output.txt
python nlp.py compress -f input.txt -l es # specify language
MLM-based compression (20-30% reduction, free, offline, predictability-aware):
python mlm.py compress "Your verbose text here"
python mlm.py compress -f input.txt -o output.txt
python mlm.py compress -f input.txt -k 30 # adjust compression level
LLM-based compression (40-58% reduction, requires API key):
python codragraph_compress.py compress "Your verbose text here"
python codragraph_compress.py compress -f input.txt -o output.txt
Decompress:
python codragraph_compress.py decompress "Codragraph text here"
| Normal (201 tokens) | Codragraph (156 tokens) |
|
I am John Smith, a 32-year-old Senior Software Engineer at a large enterprise software company based in San Francisco, California. I have over 8 years of experience in backend development, distributed systems, and database optimization. Throughout my career, I have successfully designed and implemented scalable microservices... |
John Smith. 32 years old. Senior Software Engineer. Large enterprise software company. San Francisco, California. 8 years experience. Backend development, distributed systems, database optimization. Designed scalable microservices. 50 million requests daily... |
| 22% reduction |
| Normal (171 tokens) | Codragraph (72 tokens) |
|
You are a helpful AI assistant designed to provide accurate and concise responses to user queries. When answering questions, you should always prioritize clarity and correctness over speed. If you are uncertain about any information, you must explicitly state your uncertainty... |
Helpful AI assistant. Provide accurate, concise responses. Prioritize clarity, correctness. If uncertain, state uncertainty. Break complex problems into smaller steps. Explain reasoning clearly... |
| 58% reduction |
| Normal (137 tokens) | Codragraph (79 tokens) |
|
To authenticate with our API, you need to include your API key in the Authorization header of every request. The API key should be prefixed with the word "Bearer" followed by a space. If authentication fails, the server will return a 401 Unauthorized status code... |
Authenticate API. Include API key in Authorization header every request. Prefix API key with "Bearer" space. Authentication fail, server return 401 Unauthorized status code, error message explain fail... |
| 42% reduction |
Automated benchmark verifying that specific facts are preserved and retrievable after compression:
# LLM-based compression
python benchmark/factual_preservation/llm.py
# NLP-based compression
python benchmark/factual_preservation/nlp.py
Results: 13/13 facts preserved (100%) with 12-25% compression ratio.
See benchmark/factual_preservation/ for details.
| Test Case | Original | Compressed | Reduction |
|---|---|---|---|
| System prompt | 171 tokens | 72 tokens | 58% |
| API documentation | 137 tokens | 79 tokens | 42% |
| Resume | 201 tokens | 156 tokens | 22% |
| Average | 170 | 102 | 40% |
All examples validated with GPT-4o. See examples/ for full text.
See SPEC.md for full rules.
Original:
A network router is a device that forwards data packets between computer networks. Routers perform the traffic directing functions on the Internet. When a data packet arrives at a router, the router examines the destination IP address...
Compressed:
Network router forwards data packets. Routers direct Internet traffic. Packet arrives router. Router examines destination IP address. Router determines best path. Router uses routing table...
Why it works: Store compressed docs in vector DB. Agent receives compressed RAG results directly. No decompression needed—agent understands codragraph format. Fits 2-3x more context.
Original:
First, I need to understand what the user is asking for. They want to calculate the optimal route between two cities considering both distance and traffic conditions. Let me break this down into steps. Step one: I should identify the starting city...
Compressed:
Need understand user request. User wants optimal route between cities. Consider distance, traffic. Step one: Identify starting city, destination city. Step two: Retrieve current traffic data for routes...
Why it works: Agent thinks in codragraph format during problem-solving. Chain-of-thought uses 50% fewer tokens. More reasoning steps fit in context window.
codragraph_compress.py)mlm.py)nlp.py)✅ Good for:
❌ Avoid for:
Contributions welcome. Submit issues or PRs.
MIT
William Peltomäki
Inspired by TOON and the token-optimization movement.
FAQs
Lossless semantic compression for LLM contexts — strip grammar, keep facts, save tokens.
The npm package @codragraph/compress receives a total of 10 weekly downloads. As such, @codragraph/compress popularity was classified as not popular.
We found that @codragraph/compress demonstrated a healthy version release cadence and project activity because the last version was released less than a year ago. It has 1 open source maintainer collaborating on the project.

Security News
arXiv now limits authors to two submissions a month as AI slop overwhelms moderators, delays good papers, and sparks debate over applying the limit to everyone.

Research
/Security News
A new GhostAction wave hits hundreds of GitHub repos, expanding CI/CD secret theft to cloud and AI credentials in source code and git history.

Research
/Security News
Tensorlake npm SDK version 0.5.144 was compromised in a ChainDrop / Shai-Hulud attack, delivering credential-stealing malware.