How to quickly extract text from PDF files using PHP and Alibaba Cloud OCR?-PHP Tutorial-php.cn

Home

Backend Development

PHP Tutorial

How to quickly extract text from PDF files using PHP and Alibaba Cloud OCR?

王林

Jul 19, 2023 pm 05:12 PM

php ocr pdf extraction

How to quickly extract text from PDF files using PHP and Alibaba Cloud OCR?

Introduction:
With the advent of the digital age, more and more documents are saved in PDF format. In some scenarios, we need to extract text from PDF files for further processing and analysis, such as automated document processing, information extraction, etc. This article will introduce how to use PHP and Alibaba Cloud OCR service to quickly extract text from PDF files.

Step 1: Configure Alibaba Cloud OCR service
First, we need to register and activate the OCR service on Alibaba Cloud. Obtain the Access Key ID and Access Key Secret, and create an OCR application to generate a key under the application. This information will be used in subsequent code.

Step 2: Install and configure PHP-SDK
Alibaba Cloud provides a PHP version of the SDK. We can use composer to quickly install and configure the SDK. Execute the following command in the terminal:

composer require alibabacloud/ocr-sdk-php

After the installation is complete, add the following code to the project, introduce the SDK, and configure the Access Key ID and Access Key Secret:

<?php
use AlibabaCloudClientAlibabaCloud;
use AlibabaCloudClientExceptionClientException;
use AlibabaCloudClientExceptionServerException;

AlibabaCloud::accessKeyClient('your-access-key-id', 'your-access-key-secret')
            ->regionId('cn-shanghai')
            ->asDefaultClient();
?>

Place the above code in " Replace "your-access-key-id" and "your-access-key-secret" with your actual information.

Step 3: Use the OCR service to extract PDF text
In the PHP script, we can use the "ocr_document_recognize" interface provided by Alibaba Cloud OCR to identify the PDF file and obtain the text in it.

The following is a sample code:

try {
    $result = AlibabaCloud::rpc()
              ->product('ocr')
              ->scheme('https')
              ->version('2019-12-30')
              ->action('ocr_document_recognize')
              ->method('POST')
              ->host('ocr.cn-shanghai.aliyuncs.com')
              ->options([
                'query' => [
                  'RegionId' => 'cn-shanghai',
                  'AccessKeyId' => 'your-access-key-id',
                  'AccessKeySecret' => 'your-access-key-secret',
                ],
              ])
              ->request();
    
    // 解析返回结果
    $text = '';
    foreach ($result['Data']['Regions'] as $region) {
        foreach ($region['Lines'] as $line) {
            $text .= $line['Text'] . "
";
        }
    }
    
    // 打印提取的文字
    echo $text;

} catch (ClientException $e) {
    echo $e->getErrorMessage() . PHP_EOL;
} catch (ServerException $e) {
    echo $e->getErrorMessage() . PHP_EOL;
}

Replace "your-access-key-id" and "your-access-key-secret" in the above code with your actual information.

Through the above steps, we can use PHP and Alibaba Cloud OCR service to quickly extract text from PDF files. You can further process and analyze the extracted text according to actual needs.

Summary:
This article introduces how to use PHP and Alibaba Cloud OCR service to quickly extract text from PDF files. By configuring the Alibaba Cloud OCR service and installing PHP-SDK, we can use the interface provided by Alibaba Cloud OCR to identify PDF files and extract text information in them. In this way, we can easily perform automated document processing and information extraction operations to improve work efficiency.

The above is the detailed content of How to quickly extract text from PDF files using PHP and Alibaba Cloud OCR?. For more information, please follow other related articles on the PHP Chinese website!

Statement of this Website

The content of this article is voluntarily contributed by netizens, and the copyright belongs to the original author. This site does not assume corresponding legal responsibility. If you find any content suspected of plagiarism or infringement, please contact admin@php.cn

Hot AI Tools

Undress AI Tool

Undress images for free

Undresser.AI Undress

AI-powered app for creating realistic nude photos

AI Clothes Remover

Online AI tool for removing clothes from photos.

Clothoff.io

AI clothes remover

Video Face Swap

Swap faces in any video effortlessly with our completely free AI face swap tool!

Hot Article

How to report an impersonation account on Instagram

3 weeks ago By 下次还敢

Wuchang: Fallen Feathers - Dragon Emperor Zhu Youjian Boss Fight Guide

4 weeks ago By DDD

How to Change ChatGPT Personality in Settings (Cynic, Robot, Listener, Nerd)

3 weeks ago By DDD

How to Fight Eris in Neon Abyss

3 weeks ago By Jack chen

Pokémon TCG Scarlet & Violet: Black Bolt Elite Trainer Box Review

4 weeks ago By Jack chen

Hot Tools

Notepad++7.3.1

Easy-to-use and free code editor

SublimeText3 Chinese version

Chinese version, very easy to use

Zend Studio 13.0.1

Powerful PHP integrated development environment

Dreamweaver CS6

Visual web development tools

SublimeText3 Mac version

God-level code editing software (SublimeText3)

Hot Topics

PHP Tutorial

1600

276

Related knowledge

How to work with arrays in php Aug 20, 2025 pm 07:01 PM

PHParrayshandledatacollectionsefficientlyusingindexedorassociativestructures;theyarecreatedwitharray()or[],accessedviakeys,modifiedbyassignment,iteratedwithforeach,andmanipulatedusingfunctionslikecount(),in_array(),array_key_exists(),array_push(),arr

How to use the $_COOKIE variable in php Aug 20, 2025 pm 07:00 PM

$_COOKIEisaPHPsuperglobalforaccessingcookiessentbythebrowser;cookiesaresetusingsetcookie()beforeoutput,readvia$_COOKIE['name'],updatedbyresendingwithnewvalues,anddeletedbysettinganexpiredtimestamp,withsecuritybestpracticesincludinghttponly,secureflag

Describe the Observer design pattern and its implementation in PHP. Aug 15, 2025 pm 01:54 PM

TheObserverdesignpatternenablesautomaticnotificationofdependentobjectswhenasubject'sstatechanges.1)Itdefinesaone-to-manydependencybetweenobjects;2)Thesubjectmaintainsalistofobserversandnotifiesthemviaacommoninterface;3)Observersimplementanupdatemetho

phpMyAdmin security best practices Aug 17, 2025 am 01:56 AM

To effectively protect phpMyAdmin, multiple layers of security measures must be taken. 1. Restrict access through IP, only trusted IP connections are allowed; 2. Modify the default URL path to a name that is not easy to guess; 3. Use strong passwords and create a dedicated MySQL user with minimized permissions, and it is recommended to enable two-factor authentication; 4. Keep the phpMyAdmin version up to fix known vulnerabilities; 5. Strengthen the web server and PHP configuration, disable dangerous functions and restrict file execution; 6. Force HTTPS to encrypt communication to prevent credential leakage; 7. Disable phpMyAdmin when not in use or increase HTTP basic authentication; 8. Regularly monitor logs and configure fail2ban to defend against brute force cracking; 9. Delete setup and

Using XSLT Parameters to Create Dynamic Transformations Aug 17, 2025 am 09:16 AM

XSLT parameters are a key mechanism for dynamic conversion through external passing values. 1. Use declared parameters and set default values; 2. Pass the actual value from application code (such as C#) through interfaces such as XsltArgumentList; 3. Control conditional processing, localization, data filtering or output format through $paramName reference parameters in the template; 4. Best practices include using meaningful names, providing default values, grouping related parameters, and performing value verification. The rational use of parameters can make XSLT style sheets highly reusable and maintainable, and the same style sheets can produce diversified output results based on different inputs.

You are not currently using a display attached to an NVIDIA GPU [Fixed] Aug 19, 2025 am 12:12 AM

Ifyousee"YouarenotusingadisplayattachedtoanNVIDIAGPU,"ensureyourmonitorisconnectedtotheNVIDIAGPUport,configuredisplaysettingsinNVIDIAControlPanel,updatedriversusingDDUandcleaninstall,andsettheprimaryGPUtodiscreteinBIOS/UEFI.Restartaftereach

How would you implement API versioning in a PHP application? Aug 14, 2025 pm 11:14 PM

APIversioninginPHPcanbeeffectivelyimplementedusingURL,header,orqueryparameterapproaches,withURLandheaderversioningbeingmostrecommended.1.ForURL-basedversioning,includetheversionintheroute(e.g.,/v1/users)andorganizecontrollersinversioneddirectories,ro

How to work with dates and times in php Aug 20, 2025 pm 06:57 PM

UseDateTimefordatesinPHP:createwithnewDateTime(),formatwithformat(),modifyviaadd()ormodify(),settimezoneswithDateTimeZone,andcompareusingoperatorsordiff()togetintervals.

See all articles