Deepcrawl SDK
TypeScript SDK for the Deepcrawl API - Web scraping and crawling with comprehensive error handling.

⚡ Why Deepcrawl SDK?
- 🏗️ oRPC-Powered: Built on oRPC framework for type-safe RPC
- 🔒 Type-Safe: End-to-end TypeScript with error handling
- 🖥️ Server-Side Only: Designed for Node.js, Cloudflare Workers, and Next.js Server Actions
- 🪶 Lightweight: Minimal bundle size with tree-shaking support
- 🛡️ Error Handling: Comprehensive, typed errors with context
- 🔄 Retry Logic: Built-in exponential backoff for transient failures
- ⚡ Connection Pooling: Automatic HTTP connection reuse (Node.js)
📦 Installation
npm install deepcrawl
yarn add deepcrawl
pnpm add deepcrawl
🚀 Quick Start
import { DeepcrawlApp } from 'deepcrawl';
const deepcrawl = new DeepcrawlApp({
apiKey: "dc-YOUR_API_KEY"
});
try {
const result = await deepcrawl.readUrl('https://example.com');
console.log(result.markdown);
} catch (error) {
console.error('Scraping failed:', error.message);
}
📖 API Methods
readUrl(url, options?)
Extract clean content and metadata from any URL.
const result = await deepcrawl.readUrl('https://example.com', {
metadata: true,
markdown: true,
cleanedHtml: true,
rawHtml: true,
metricsOptions: {
enable: true
}
});
interface ReadUrlResponse {
targetUrl: string;
success: boolean;
markdown?: string;
cleanedHtml?: string;
rawHtml?: string;
metadata?: {
title?: string;
description?: string;
author?: string;
publishedTime?: string;
ogImage?: string;
favicon?: string;
};
metrics?: {
readableDuration: string;
durationMs: number;
startTimeMs: number;
endTimeMs: number;
};
}
getMarkdown(url, options?)
Simplified method to get just markdown content.
const result = await deepcrawl.getMarkdown('https://example.com', {
metricsOptions: { enable: true }
});
interface GetMarkdownResponse {
targetUrl: string;
success: boolean;
markdown: string;
metrics?: {
readableDuration: string;
durationMs: number;
startTimeMs: number;
endTimeMs: number;
};
}
Extract all links from a page with powerful filtering options.
const result = await deepcrawl.extractLinks('https://example.com', {
includeInternal: true,
includeExternal: false,
includeEmails: false,
includePhoneNumbers: false,
includeSocialMedia: false,
metricsOptions: { enable: true }
});
interface ExtractLinksResponse {
targetUrl: string;
success: boolean;
tree: {
internal: Array<{ href: string; text: string }>;
external: Array<{ href: string; text: string }>;
emails: string[];
phoneNumbers: string[];
socialMedia: Array<{ platform: string; url: string }>;
};
metrics?: {
readableDuration: string;
durationMs: number;
startTimeMs: number;
endTimeMs: number;
};
}
getManyLogs(options?)
Retrieve activity logs with paginated results and filtering by path, success status, and date range. Returns logs with full type safety through discriminated unions based on the endpoint path.
const result = await deepcrawl.getManyLogs({
limit: 50,
offset: 0,
path: 'read-getMarkdown',
success: true,
startDate: '2025-01-01T00:00:00Z',
endDate: '2025-12-31T23:59:59Z',
orderBy: 'requestTimestamp',
orderDir: 'desc'
});
interface GetManyLogsResponse {
logs: ActivityLogEntry[];
meta: {
limit: number;
offset: number;
hasMore: boolean;
nextOffset: number | null;
orderBy: string;
orderDir: 'asc' | 'desc';
startDate?: string;
endDate?: string;
};
}
getOneLog(options)
Get a single activity log entry by ID with full type safety through discriminated unions.
const log = await deepcrawl.getOneLog({
id: 'request-id-123'
});
🌟 Real-World Usage Examples
1. E-commerce Product Monitoring
import { DeepcrawlApp, DeepcrawlError } from 'deepcrawl';
async function monitorProduct(productUrl: string) {
const deepcrawl = new DeepcrawlApp({ apiKey: process.env.DEEPCRAWL_API_KEY! });
try {
const result = await deepcrawl.readUrl(productUrl, {
metadata: true,
cleanedHtml: true
});
return {
title: result.metadata?.title,
price: extractPrice(result.cleanedHtml),
availability: checkAvailability(result.cleanedHtml),
lastChecked: new Date().toISOString()
};
} catch (error) {
if (DeepcrawlError.isRateLimitError(error)) {
console.log(`Rate limited. Retrying in ${error.retryAfter} seconds...`);
await delay(error.retryAfter * 1000);
return monitorProduct(productUrl);
}
if (DeepcrawlError.isReadError(error)) {
throw new Error(`Failed to scrape ${error.targetUrl}: ${error.userMessage}`);
}
throw error;
}
}
2. Content Aggregation Pipeline
import { DeepcrawlApp, DeepcrawlError } from 'deepcrawl';
class ContentAggregator {
private deepcrawl = new DeepcrawlApp({ apiKey: process.env.DEEPCRAWL_API_KEY! });
private readonly maxRetries = 3;
async aggregateArticles(urls: string[]) {
const results = await Promise.allSettled(
urls.map(url => this.scrapeWithRetry(url))
);
return results.map((result, index) => ({
url: urls[index],
success: result.status === 'fulfilled',
data: result.status === 'fulfilled' ? result.value : null,
error: result.status === 'rejected' ? result.reason.message : null
}));
}
private async scrapeWithRetry(url: string, attempt = 1): Promise<Article> {
try {
const result = await this.deepcrawl.readUrl(url, {
metadata: true,
markdown: true
});
return {
title: result.metadata?.title || 'Untitled',
content: result.markdown,
publishedAt: result.metadata?.publishedTime,
author: result.metadata?.author,
sourceUrl: url
};
} catch (error) {
if (error instanceof DeepcrawlError) {
if (error.isRateLimit() && attempt <= this.maxRetries) {
await delay(error.retryAfter * 1000 * attempt);
return this.scrapeWithRetry(url, attempt + 1);
}
if (error.isNetwork() && attempt <= this.maxRetries) {
await delay(1000 * attempt);
return this.scrapeWithRetry(url, attempt + 1);
}
if (error.isRead()) {
throw new Error(`Content unavailable: ${error.userMessage}`);
}
}
throw error;
}
}
}
3. Next.js Server Actions with Rich Error Handling
'use server'
import { DeepcrawlApp, DeepcrawlError } from 'deepcrawl';
import { headers } from 'next/headers';
import { revalidatePath } from 'next/cache';
export async function scrapeUrlAction(url: string) {
const deepcrawl = new DeepcrawlApp({
apiKey: process.env.DEEPCRAWL_API_KEY!,
headers: await headers(),
});
try {
const result = await deepcrawl.readUrl(url, {
metadata: true,
markdown: true,
});
await saveToDatabase(url, result);
revalidatePath('/dashboard');
return {
success: true,
data: {
title: result.metadata?.title,
description: result.metadata?.description,
content: result.markdown,
targetUrl: result.targetUrl
}
};
} catch (error) {
if (error instanceof DeepcrawlError) {
return {
success: false,
error: {
type: error.constructor.name,
message: error.userMessage,
retryable: error.isRateLimit() || error.isNetwork(),
retryAfter: error.isRateLimit() ? error.retryAfter : undefined
}
};
}
return {
success: false,
error: {
type: 'UnknownError',
message: 'An unexpected error occurred',
retryable: false
}
};
}
}
4. React Hook with Comprehensive Error States
import { useState, useCallback } from 'react';
import { DeepcrawlApp, DeepcrawlError } from 'deepcrawl';
interface UseScrapingState {
data: any | null;
loading: boolean;
error: string | null;
retryInfo: { canRetry: boolean; retryAfter?: number } | null;
}
export function useScraping(apiKey: string) {
const [state, setState] = useState<UseScrapingState>({
data: null,
loading: false,
error: null,
retryInfo: null
});
const deepcrawl = new DeepcrawlApp({ apiKey });
const scrape = useCallback(async (url: string) => {
setState(prev => ({ ...prev, loading: true, error: null, retryInfo: null }));
try {
const result = await deepcrawl.readUrl(url, { metadata: true });
setState({
data: result,
loading: false,
error: null,
retryInfo: null
});
} catch (error) {
if (error instanceof DeepcrawlError) {
setState({
data: null,
loading: false,
error: error.userMessage,
retryInfo: {
canRetry: error.isRateLimit() || error.isNetwork(),
retryAfter: error.isRateLimit() ? error.retryAfter : undefined
}
});
} else {
setState({
data: null,
loading: false,
error: 'An unexpected error occurred',
retryInfo: { canRetry: false }
});
}
}
}, [deepcrawl]);
const retry = useCallback(() => {
if (state.retryInfo?.canRetry) {
scrape(state.data?.targetUrl || '');
}
}, [state.retryInfo, state.data, scrape]);
return { ...state, scrape, retry };
}
5. Activity Logging with Server Actions
'use server';
import { deepcrawlClient } from '@/lib/deepcrawl';
export async function fetchDeepcrawlLogs(filters?: {
userId?: string;
url?: string;
operation?: 'readUrl' | 'getMarkdown' | 'extractLinks';
status?: 'success' | 'error';
limit?: number;
offset?: number;
}) {
return deepcrawlClient.getManyLogs(filters);
}
'use server';
import { deepcrawlClient } from '@/lib/deepcrawl';
export async function getActivityLogs() {
try {
const logs = await deepcrawlClient.getManyLogs({
limit: 50,
offset: 0
});
return { success: true, data: logs };
} catch (error) {
return {
success: false,
error: error instanceof Error ? error.message : 'Failed to fetch logs'
};
}
}
'use client';
import { useState, useEffect } from 'react';
import { getActivityLogs } from '@/app/actions/logs';
export function ActivityLogsClient() {
const [logs, setLogs] = useState([]);
const [loading, setLoading] = useState(true);
useEffect(() => {
getActivityLogs().then(result => {
if (result.success) {
setLogs(result.data);
}
setLoading(false);
});
}, []);
if (loading) return <div>Loading...</div>;
return (
<div>
{logs.map(log => (
<div key={log.id}>
<p>{log.operation} - {log.url}</p>
<p>Status: {log.status}</p>
{log.errorMessage && <p>Error: {log.errorMessage}</p>}
</div>
))}
</div>
);
}
🛡️ Error Handling Patterns
The SDK provides multiple patterns for different coding styles:
Traditional Try/Catch
try {
const result = await deepcrawl.readUrl(url);
} catch (error) {
if (error instanceof DeepcrawlReadError) {
console.log(`Failed to read ${error.targetUrl}: ${error.message}`);
}
}
Static Type Guards
const [error, result] = await safe(deepcrawl.readUrl(url));
if (DeepcrawlError.isRateLimitError(error)) {
await delay(error.retryAfter * 1000);
}
Instance Methods
try {
const result = await deepcrawl.readUrl(url);
} catch (error) {
if (error.isRateLimit?.()) {
console.log(`Retry after ${error.retryAfter}s`);
}
}
📚 Error Types Reference
Business Logic Errors
DeepcrawlReadError - Content extraction failed
DeepcrawlLinksError - Link extraction failed
Infrastructure Errors
DeepcrawlRateLimitError - Rate limit exceeded (includes retryAfter)
DeepcrawlAuthError - Authentication failed
DeepcrawlValidationError - Invalid request parameters
DeepcrawlNotFoundError - Resource not found
DeepcrawlServerError - Server-side error
DeepcrawlNetworkError - Network connectivity issues
Rich Error Properties
interface ErrorProperties {
code: string;
status: number;
message: string;
data: any;
defined: boolean;
userMessage: string;
targetUrl: string;
success: false;
error: string;
retryAfter: number;
operation: string;
timestamp: string;
tree?: any;
}
🔧 Configuration
const deepcrawl = new DeepcrawlApp({
apiKey: "dc-YOUR_API_KEY",
baseUrl: "https://api.deepcrawl.dev",
headers: {
'User-Agent': 'MyApp/1.0'
},
fetch: customFetch,
fetchOptions: {
timeout: 30000
}
});
Connection Pooling (Node.js Only)
The SDK automatically uses HTTP connection pooling in Node.js environments:
{
keepAlive: true,
maxSockets: 10,
maxFreeSockets: 5,
timeout: 60000,
keepAliveMsecs: 30000
}
Benefits:
- ⚡ Faster for concurrent requests
- 🔄 Connection reuse reduces handshake overhead
- 🎯 Auto-cleanup of idle connections
- 📊 Optimized for batch operations
Note: Connection pooling is automatically disabled in browser and Edge Runtime environments.
🔄 Built-in Retry Logic
The SDK includes automatic retry logic with exponential backoff for transient failures:
const result = await deepcrawl.readUrl('https://example.com');
Retry behavior:
- Rate Limits: Waits for
retryAfter seconds before retry
- Network Errors: Exponential backoff (1s, 2s, 4s)
- Max Attempts: 3 retries by default
- Customizable: Handle errors manually for custom retry logic
🔒 Security Best Practices
Next.js Server Actions (Recommended)
Always use Server Actions to keep your API key secure:
'use server';
export const DEEPCRAWL_API_KEY = process.env.DEEPCRAWL_API_KEY as string;
export const deepcrawlClient = new DeepcrawlApp({ apiKey: DEEPCRAWL_API_KEY });
'use server';
import { deepcrawlClient } from '@/lib/deepcrawl';
export async function scrapeAction(url: string) {
return deepcrawlClient.readUrl(url);
}
'use client';
import { scrapeAction } from '@/app/actions/scrape';
export function ScrapeButton() {
const handleClick = async () => {
const result = await scrapeAction('https://example.com');
console.log(result);
};
return <button onClick={handleClick}>Scrape</button>;
}
What NOT to Do
'use client';
import { DeepcrawlApp } from 'deepcrawl';
export function BadComponent() {
const deepcrawl = new DeepcrawlApp({
apiKey: process.env.DEEPCRAWL_API_KEY
});
}
🌍 Environment Support
⚠️ Server-Side Only: The Deepcrawl SDK requires an API key and is designed for server-side use only:
- ✅ Node.js (18+) - with connection pooling
- ✅ Cloudflare Workers
- ✅ Vercel Edge Runtime
- ✅ Next.js Server Actions (recommended)
- ✅ Deno, Bun, and other modern runtimes
- ❌ Browser environments (use Server Actions instead)
Why Server-Side Only?
- API keys must remain secret and never be exposed to client-side code
- For client-side functionality, use Next.js Server Actions as shown in the examples above
- This architecture ensures your API key stays secure while still enabling client-side interactions
Runtime detection:
const runtime = deepcrawl.nodeEnv;
📦 TypeScript Types
All types are fully exported:
import type {
DeepcrawlApp,
DeepcrawlConfig,
ReadUrlOptions,
ReadUrlResponse,
GetMarkdownOptions,
GetMarkdownResponse,
ExtractLinksOptions,
ExtractLinksResponse,
ActivityLog,
ActivityLogFilters,
DeepcrawlError,
DeepcrawlReadError,
DeepcrawlLinksError,
DeepcrawlRateLimitError,
DeepcrawlAuthError,
DeepcrawlValidationError,
DeepcrawlNotFoundError,
DeepcrawlServerError,
DeepcrawlNetworkError,
Metadata,
MetricsOptions,
Metrics,
} from 'deepcrawl';
📄 License
MIT - see LICENSE for details.
🤝 Support
Built with ❤️ by the Deepcrawl team