New:Microsoft Teams Notifications Are Now Available in Socket.Learn more →
Get Started

deepcrawl

Package Overview
Dependencies
Maintainers
1
Versions
53
Alerts
File Explorer

Advanced tools

Socket logo

Install Socket

Detect and block malicious and high-risk dependencies

Install

deepcrawl

JavaScript/TypeScript SDK for Deepcrawl API - A powerful web scraping and crawling service

Source
npmnpm
Version
0.4.5
Version published
Weekly downloads
111
-17.78%
Maintainers
1
Weekly downloads
 
Created
Source

Deepcrawl SDK

TypeScript SDK for the Deepcrawl API - Web scraping and crawling with comprehensive error handling.

npm version TypeScript MIT License

⚡ Why Deepcrawl SDK?

  • 🏗️ oRPC-Powered: Built on oRPC framework for type-safe RPC
  • 🔒 Type-Safe: End-to-end TypeScript with error handling
  • 🖥️ Server-Side Only: Designed for Node.js, Cloudflare Workers, and Next.js Server Actions
  • 🪶 Lightweight: Minimal bundle size with tree-shaking support
  • 🛡️ Error Handling: Comprehensive, typed errors with context
  • 🔄 Retry Logic: Built-in exponential backoff for transient failures
  • ⚡ Connection Pooling: Automatic HTTP connection reuse (Node.js)

📦 Installation

npm install deepcrawl
# or
yarn add deepcrawl
# or
pnpm add deepcrawl

🚀 Quick Start

import { DeepcrawlApp } from 'deepcrawl';

const deepcrawl = new DeepcrawlApp({
  apiKey: "dc-YOUR_API_KEY"
});

try {
  const result = await deepcrawl.readUrl('https://example.com');
  console.log(result.markdown);
} catch (error) {
  console.error('Scraping failed:', error.message);
}

📖 API Methods

readUrl(url, options?)

Extract clean content and metadata from any URL.

const result = await deepcrawl.readUrl('https://example.com', {
  metadata: true,        // Extract page metadata
  markdown: true,        // Convert to markdown
  cleanedHtml: true,     // Get sanitized HTML
  rawHtml: true,        // Get original HTML
  metricsOptions: {     // Performance tracking
    enable: true
  }
});

// Response type
interface ReadUrlResponse {
  targetUrl: string;
  success: boolean;
  markdown?: string;
  cleanedHtml?: string;
  rawHtml?: string;
  metadata?: {
    title?: string;
    description?: string;
    author?: string;
    publishedTime?: string;
    ogImage?: string;
    favicon?: string;
  };
  metrics?: {
    readableDuration: string;
    durationMs: number;
    startTimeMs: number;
    endTimeMs: number;
  };
}

getMarkdown(url, options?)

Simplified method to get just markdown content.

const result = await deepcrawl.getMarkdown('https://example.com', {
  metricsOptions: { enable: true }
});

// Response type
interface GetMarkdownResponse {
  targetUrl: string;
  success: boolean;
  markdown: string;
  metrics?: {
    readableDuration: string;
    durationMs: number;
    startTimeMs: number;
    endTimeMs: number;
  };
}

extractLinks(url, options?)

Extract all links from a page with powerful filtering options.

const result = await deepcrawl.extractLinks('https://example.com', {
  includeInternal: true,
  includeExternal: false,
  includeEmails: false,
  includePhoneNumbers: false,
  includeSocialMedia: false,
  metricsOptions: { enable: true }
});

// Response type
interface ExtractLinksResponse {
  targetUrl: string;
  success: boolean;
  tree: {
    internal: Array<{ href: string; text: string }>;
    external: Array<{ href: string; text: string }>;
    emails: string[];
    phoneNumbers: string[];
    socialMedia: Array<{ platform: string; url: string }>;
  };
  metrics?: {
    readableDuration: string;
    durationMs: number;
    startTimeMs: number;
    endTimeMs: number;
  };
}

getManyLogs(options?)

Retrieve activity logs with paginated results and filtering by path, success status, and date range. Returns logs with full type safety through discriminated unions based on the endpoint path.

const result = await deepcrawl.getManyLogs({
  limit: 50,                          // Max results (default: 20, max: 100)
  offset: 0,                          // Skip first N results (default: 0)
  path: 'read-getMarkdown',           // Filter by endpoint path (optional)
  success: true,                      // Filter by success status (optional)
  startDate: '2025-01-01T00:00:00Z',  // Filter from date (ISO 8601) (optional)
  endDate: '2025-12-31T23:59:59Z',    // Filter to date (ISO 8601) (optional)
  orderBy: 'requestTimestamp',        // Sort column (default: 'requestTimestamp')
  orderDir: 'desc'                    // Sort direction: 'asc' | 'desc' (default: 'desc')
});

// Response type with discriminated unions
interface GetManyLogsResponse {
  logs: ActivityLogEntry[];  // Array of log entries with discriminated unions
  meta: {
    limit: number;           // Effective limit applied
    offset: number;          // Effective offset applied
    hasMore: boolean;        // More logs available?
    nextOffset: number | null; // Next page offset (null if no more data)
    orderBy: string;         // Column used for sorting
    orderDir: 'asc' | 'desc'; // Sort direction applied
    startDate?: string;      // Normalized start date boundary
    endDate?: string;        // Normalized end date boundary
  };
}

// Each log entry uses discriminated union based on 'path' field:
// - 'read-getMarkdown': response is string
// - 'read-readUrl': response is ReadSuccessResponse | ReadErrorResponse
// - 'links-getLinks': response is LinksSuccessResponse | LinksErrorResponse
// - 'links-extractLinks': response is LinksSuccessResponse | LinksErrorResponse

getOneLog(options)

Get a single activity log entry by ID with full type safety through discriminated unions.

const log = await deepcrawl.getOneLog({
  id: 'request-id-123'  // Request ID (required)
});

// Response is ActivityLogEntry with discriminated union
// TypeScript automatically narrows types based on log.path

🌟 Real-World Usage Examples

1. E-commerce Product Monitoring

import { DeepcrawlApp, DeepcrawlError } from 'deepcrawl';

async function monitorProduct(productUrl: string) {
  const deepcrawl = new DeepcrawlApp({ apiKey: process.env.DEEPCRAWL_API_KEY! });

  try {
    const result = await deepcrawl.readUrl(productUrl, {
      metadata: true,
      cleanedHtml: true
    });

    return {
      title: result.metadata?.title,
      price: extractPrice(result.cleanedHtml),
      availability: checkAvailability(result.cleanedHtml),
      lastChecked: new Date().toISOString()
    };
  } catch (error) {
    if (DeepcrawlError.isRateLimitError(error)) {
      // Smart retry with exponential backoff
      console.log(`Rate limited. Retrying in ${error.retryAfter} seconds...`);
      await delay(error.retryAfter * 1000);
      return monitorProduct(productUrl);
    }

    if (DeepcrawlError.isReadError(error)) {
      throw new Error(`Failed to scrape ${error.targetUrl}: ${error.userMessage}`);
    }

    throw error;
  }
}

2. Content Aggregation Pipeline

import { DeepcrawlApp, DeepcrawlError } from 'deepcrawl';

class ContentAggregator {
  private deepcrawl = new DeepcrawlApp({ apiKey: process.env.DEEPCRAWL_API_KEY! });
  private readonly maxRetries = 3;

  async aggregateArticles(urls: string[]) {
    const results = await Promise.allSettled(
      urls.map(url => this.scrapeWithRetry(url))
    );

    return results.map((result, index) => ({
      url: urls[index],
      success: result.status === 'fulfilled',
      data: result.status === 'fulfilled' ? result.value : null,
      error: result.status === 'rejected' ? result.reason.message : null
    }));
  }

  private async scrapeWithRetry(url: string, attempt = 1): Promise<Article> {
    try {
      const result = await this.deepcrawl.readUrl(url, {
        metadata: true,
        markdown: true
      });

      return {
        title: result.metadata?.title || 'Untitled',
        content: result.markdown,
        publishedAt: result.metadata?.publishedTime,
        author: result.metadata?.author,
        sourceUrl: url
      };
    } catch (error) {
      if (error instanceof DeepcrawlError) {
        // Use instance methods for fluent checking
        if (error.isRateLimit() && attempt <= this.maxRetries) {
          await delay(error.retryAfter * 1000 * attempt); // Exponential backoff
          return this.scrapeWithRetry(url, attempt + 1);
        }

        if (error.isNetwork() && attempt <= this.maxRetries) {
          await delay(1000 * attempt);
          return this.scrapeWithRetry(url, attempt + 1);
        }

        if (error.isRead()) {
          throw new Error(`Content unavailable: ${error.userMessage}`);
        }
      }

      throw error;
    }
  }
}

3. Next.js Server Actions with Rich Error Handling

// app/actions/scrape.ts
'use server'

import { DeepcrawlApp, DeepcrawlError } from 'deepcrawl';
import { headers } from 'next/headers';
import { revalidatePath } from 'next/cache';

export async function scrapeUrlAction(url: string) {
  const deepcrawl = new DeepcrawlApp({
    apiKey: process.env.DEEPCRAWL_API_KEY!,
    headers: await headers(), // Automatic session forwarding
  });

  try {
    const result = await deepcrawl.readUrl(url, {
      metadata: true,
      markdown: true,
    });

    // Cache the result
    await saveToDatabase(url, result);
    revalidatePath('/dashboard');

    return {
      success: true,
      data: {
        title: result.metadata?.title,
        description: result.metadata?.description,
        content: result.markdown,
        targetUrl: result.targetUrl
      }
    };
  } catch (error) {
    if (error instanceof DeepcrawlError) {
      return {
        success: false,
        error: {
          type: error.constructor.name,
          message: error.userMessage,
          retryable: error.isRateLimit() || error.isNetwork(),
          retryAfter: error.isRateLimit() ? error.retryAfter : undefined
        }
      };
    }

    return {
      success: false,
      error: {
        type: 'UnknownError',
        message: 'An unexpected error occurred',
        retryable: false
      }
    };
  }
}

4. React Hook with Comprehensive Error States

import { useState, useCallback } from 'react';
import { DeepcrawlApp, DeepcrawlError } from 'deepcrawl';

interface UseScrapingState {
  data: any | null;
  loading: boolean;
  error: string | null;
  retryInfo: { canRetry: boolean; retryAfter?: number } | null;
}

export function useScraping(apiKey: string) {
  const [state, setState] = useState<UseScrapingState>({
    data: null,
    loading: false,
    error: null,
    retryInfo: null
  });

  const deepcrawl = new DeepcrawlApp({ apiKey });

  const scrape = useCallback(async (url: string) => {
    setState(prev => ({ ...prev, loading: true, error: null, retryInfo: null }));

    try {
      const result = await deepcrawl.readUrl(url, { metadata: true });
      setState({
        data: result,
        loading: false,
        error: null,
        retryInfo: null
      });
    } catch (error) {
      if (error instanceof DeepcrawlError) {
        setState({
          data: null,
          loading: false,
          error: error.userMessage,
          retryInfo: {
            canRetry: error.isRateLimit() || error.isNetwork(),
            retryAfter: error.isRateLimit() ? error.retryAfter : undefined
          }
        });
      } else {
        setState({
          data: null,
          loading: false,
          error: 'An unexpected error occurred',
          retryInfo: { canRetry: false }
        });
      }
    }
  }, [deepcrawl]);

  const retry = useCallback(() => {
    if (state.retryInfo?.canRetry) {
      // Re-trigger with the last URL
      scrape(state.data?.targetUrl || '');
    }
  }, [state.retryInfo, state.data, scrape]);

  return { ...state, scrape, retry };
}

5. Activity Logging with Server Actions

// app/query/logs-query.server.ts
'use server';

import { deepcrawlClient } from '@/lib/deepcrawl';

/**
 * Server Action: Fetch activity logs with type-safe filtering
 * @param filters - Optional filters for logs
 */
export async function fetchDeepcrawlLogs(filters?: {
  userId?: string;
  url?: string;
  operation?: 'readUrl' | 'getMarkdown' | 'extractLinks';
  status?: 'success' | 'error';
  limit?: number;
  offset?: number;
}) {
  return deepcrawlClient.getManyLogs(filters);
}

// app/actions/logs.ts
'use server';

import { deepcrawlClient } from '@/lib/deepcrawl';

export async function getActivityLogs() {
  try {
    const logs = await deepcrawlClient.getManyLogs({
      limit: 50,
      offset: 0
    });
    return { success: true, data: logs };
  } catch (error) {
    return {
      success: false,
      error: error instanceof Error ? error.message : 'Failed to fetch logs'
    };
  }
}

// app/components/activity-logs.client.tsx
'use client';

import { useState, useEffect } from 'react';
import { getActivityLogs } from '@/app/actions/logs';

export function ActivityLogsClient() {
  const [logs, setLogs] = useState([]);
  const [loading, setLoading] = useState(true);

  useEffect(() => {
    getActivityLogs().then(result => {
      if (result.success) {
        setLogs(result.data);
      }
      setLoading(false);
    });
  }, []);

  if (loading) return <div>Loading...</div>;

  return (
    <div>
      {logs.map(log => (
        <div key={log.id}>
          <p>{log.operation} - {log.url}</p>
          <p>Status: {log.status}</p>
          {log.errorMessage && <p>Error: {log.errorMessage}</p>}
        </div>
      ))}
    </div>
  );
}

🛡️ Error Handling Patterns

The SDK provides multiple patterns for different coding styles:

Traditional Try/Catch

try {
  const result = await deepcrawl.readUrl(url);
} catch (error) {
  if (error instanceof DeepcrawlReadError) {
    console.log(`Failed to read ${error.targetUrl}: ${error.message}`);
  }
}

Static Type Guards

const [error, result] = await safe(deepcrawl.readUrl(url));
if (DeepcrawlError.isRateLimitError(error)) {
  await delay(error.retryAfter * 1000);
}

Instance Methods

try {
  const result = await deepcrawl.readUrl(url);
} catch (error) {
  if (error.isRateLimit?.()) {
    console.log(`Retry after ${error.retryAfter}s`);
  }
}

📚 Error Types Reference

Business Logic Errors

  • DeepcrawlReadError - Content extraction failed
  • DeepcrawlLinksError - Link extraction failed

Infrastructure Errors

  • DeepcrawlRateLimitError - Rate limit exceeded (includes retryAfter)
  • DeepcrawlAuthError - Authentication failed
  • DeepcrawlValidationError - Invalid request parameters
  • DeepcrawlNotFoundError - Resource not found
  • DeepcrawlServerError - Server-side error
  • DeepcrawlNetworkError - Network connectivity issues

Rich Error Properties

interface ErrorProperties {
  // All errors (consistent across all error types)
  code: string;           // oRPC error code
  status: number;         // HTTP status
  message: string;        // The actual error message (always user-friendly)
  data: any;             // Raw error data from API
  defined: boolean;       // Whether this is a contract-defined error

  // Convenience getters (same as message for consistency)
  userMessage: string;    // Always same as message

  // Read/Links errors (typed getters for convenience)
  targetUrl: string;      // URL that failed (data.targetUrl)
  success: false;         // Always false for errors (data.success)
  error: string;          // Raw error from API (data.error, same as message)

  // Rate limit errors
  retryAfter: number;     // Seconds to wait (data.retryAfter)
  operation: string;      // What operation was rate limited (data.operation)

  // Links errors
  timestamp: string;      // When the error occurred (data.timestamp)
  tree?: any;            // Partial results if available (data.tree)
}

🔧 Configuration

const deepcrawl = new DeepcrawlApp({
  apiKey: "dc-YOUR_API_KEY",           // Required
  baseUrl: "https://api.deepcrawl.dev", // Optional (default: production URL)
  headers: {                            // Optional (can be HeadersInit or Next.js headers())
    'User-Agent': 'MyApp/1.0'
  },
  fetch: customFetch,                   // Optional (custom fetch implementation)
  fetchOptions: {                       // Optional (passed to RPCLink)
    timeout: 30000
  }
});

Connection Pooling (Node.js Only)

The SDK automatically uses HTTP connection pooling in Node.js environments:

// Automatic configuration (no action needed)
{
  keepAlive: true,        // Reuse connections
  maxSockets: 10,         // Max concurrent connections per host
  maxFreeSockets: 5,      // Keep 5 idle connections ready
  timeout: 60000,         // 60s socket timeout
  keepAliveMsecs: 30000   // Send keepalive probe every 30s
}

Benefits:

  • ⚡ Faster for concurrent requests
  • 🔄 Connection reuse reduces handshake overhead
  • 🎯 Auto-cleanup of idle connections
  • 📊 Optimized for batch operations

Note: Connection pooling is automatically disabled in browser and Edge Runtime environments.

🔄 Built-in Retry Logic

The SDK includes automatic retry logic with exponential backoff for transient failures:

// Automatic retry for rate limits and network errors
const result = await deepcrawl.readUrl('https://example.com');
// If rate limited, automatically waits and retries
// If network error, retries with exponential backoff

Retry behavior:

  • Rate Limits: Waits for retryAfter seconds before retry
  • Network Errors: Exponential backoff (1s, 2s, 4s)
  • Max Attempts: 3 retries by default
  • Customizable: Handle errors manually for custom retry logic

🔒 Security Best Practices

Always use Server Actions to keep your API key secure:

// ✅ SECURE: lib/deepcrawl.ts
'use server';

export const DEEPCRAWL_API_KEY = process.env.DEEPCRAWL_API_KEY as string;
export const deepcrawlClient = new DeepcrawlApp({ apiKey: DEEPCRAWL_API_KEY });
// ✅ SECURE: app/actions/scrape.ts
'use server';

import { deepcrawlClient } from '@/lib/deepcrawl';

export async function scrapeAction(url: string) {
  return deepcrawlClient.readUrl(url);
}
// ✅ SECURE: Client component uses Server Action
'use client';

import { scrapeAction } from '@/app/actions/scrape';

export function ScrapeButton() {
  const handleClick = async () => {
    const result = await scrapeAction('https://example.com');
    console.log(result);
  };

  return <button onClick={handleClick}>Scrape</button>;
}

What NOT to Do

// ❌ INSECURE: Direct SDK usage in client components
'use client';

import { DeepcrawlApp } from 'deepcrawl';

export function BadComponent() {
  const deepcrawl = new DeepcrawlApp({
    apiKey: process.env.DEEPCRAWL_API_KEY // ❌ Exposes API key to browser!
  });
  // ...
}

🌍 Environment Support

⚠️ Server-Side Only: The Deepcrawl SDK requires an API key and is designed for server-side use only:

  • ✅ Node.js (18+) - with connection pooling
  • ✅ Cloudflare Workers
  • ✅ Vercel Edge Runtime
  • ✅ Next.js Server Actions (recommended)
  • ✅ Deno, Bun, and other modern runtimes
  • ❌ Browser environments (use Server Actions instead)

Why Server-Side Only?

  • API keys must remain secret and never be exposed to client-side code
  • For client-side functionality, use Next.js Server Actions as shown in the examples above
  • This architecture ensures your API key stays secure while still enabling client-side interactions

Runtime detection:

// Automatic detection
const runtime = deepcrawl.nodeEnv;
// Returns: 'nodeJs' | 'browser' | 'cf-worker'

📦 TypeScript Types

All types are fully exported:

import type {
  // Client
  DeepcrawlApp,
  DeepcrawlConfig,

  // API Methods
  ReadUrlOptions,
  ReadUrlResponse,
  GetMarkdownOptions,
  GetMarkdownResponse,
  ExtractLinksOptions,
  ExtractLinksResponse,

  // Activity Logs
  ActivityLog,
  ActivityLogFilters,

  // Errors
  DeepcrawlError,
  DeepcrawlReadError,
  DeepcrawlLinksError,
  DeepcrawlRateLimitError,
  DeepcrawlAuthError,
  DeepcrawlValidationError,
  DeepcrawlNotFoundError,
  DeepcrawlServerError,
  DeepcrawlNetworkError,

  // Metadata
  Metadata,
  MetricsOptions,
  Metrics,
} from 'deepcrawl';

📄 License

MIT - see LICENSE for details.

🤝 Support

Built with ❤️ by the Deepcrawl team

Keywords

deepcrawl

FAQs

Package last updated on 14 Oct 2025

Related posts