pdftohtml.1 2.47 KB
Newer Older
1 2 3 4
.TH PDFTOHTML 1
.\" NAME should be all caps, SECTION should be 1-8, maybe w/ subsection
.\" other parms are allowed: see man(7), man(1)
.SH NAME
5
pdftohtml \- program to convert PDF files into HTML, XML and PNG images
6 7
.SH SYNOPSIS
.B pdftohtml
8
.I "[options] <PDF-file> [<HTML-file> <XML-file>]"
9 10 11 12 13 14 15 16
.SH "DESCRIPTION"
This manual page documents briefly the
.BR pdftohtml 
command.
This manual page was written for the Debian GNU/Linux distribution
because the original program does not have a manual page.
.PP
.B pdftohtml
17
is a program that converts PDF documents into HTML. It generates its output in
18 19 20 21 22 23 24 25 26 27 28 29 30 31
the current working directory.
.SH OPTIONS
A summary of options are included below.
.TP
.B \-h, \-help
Show summary of options.
.TP
.B \-f <int>
first page to print
.TP
.B \-l <int>
last page to print
.TP
.B \-q
32
do not print any messages or errors
33 34 35 36 37 38 39 40 41 42
.TP
.B \-v
print copyright and version info
.TP
.B \-p
exchange .pdf links with .html
.TP
.B \-c
generate complex output
.TP
Albert Astals Cid's avatar
Albert Astals Cid committed
43
.B \-s
44
generate single HTML that includes all pages
Albert Astals Cid's avatar
Albert Astals Cid committed
45
.TP
46 47 48 49 50 51 52 53 54 55
.B \-i
ignore images
.TP
.B \-noframes
generate no frames. Not supported in complex output mode.
.TP
.B \-stdout
use standard output
.TP 
.B \-zoom <fp>
56
zoom the PDF document (default 1.5)
57 58 59 60
.TP
.B \-xml
output for XML post-processing
.TP
61 62 63
.B \-noRoundedCoordinates
do not round coordinates (with XML output only)
.TP
64 65 66 67 68 69 70 71 72 73 74 75
.B \-enc <string>
output text encoding name
.TP
.B \-opw <string>
owner password (for encrypted files)
.TP
.B \-upw <string>
user password (for encrypted files)
.TP
.B \-hidden
force hidden text extraction
.TP
76
.B \-fmt
77
image file format for Splash output (png or jpg).
78
If complex is selected, but \-fmt is not specified,
79
\-fmt png will be assumed
80 81 82 83 84 85
.TP
.B \-nomerge
do not merge paragraphs
.TP
.B \-nodrm
override document DRM settings
86 87 88 89 90
.TP
.B \-wbt <fp>
adjust the word break threshold percent. Default is 10.
Word break occurs when distance between two adjacent characters is
greater than this percent of character height.
91 92 93
.TP
.B \-fontfullname
outputs the font name without any substitutions.
94 95 96 97 98 99

.SH AUTHOR

Pdftohtml was developed by Gueorgui Ovtcharov and Rainer Dorsch. It is
based and benefits a lot from Derek Noonburg's xpdf package.

100
This manual page was written by Søren Boll Overgaard <boll@debian.org>,
101
for the Debian GNU/Linux system (but may be used by others).
102
.SH "SEE ALSO"
103
.BR pdfdetach (1),
104 105 106 107 108 109 110
.BR pdffonts (1),
.BR pdfimages (1),
.BR pdfinfo (1),
.BR pdftocairo (1),
.BR pdftoppm (1),
.BR pdftops (1),
.BR pdftotext (1)
111 112
.BR pdfseparate (1),
.BR pdfsig (1),
113
.BR pdfunite (1)