PHP 中的 DOCX 文件类型finfo_file是应用程序/zip
2022-08-30 20:00:32
你好,我正在尝试通过finfo_file函数来验证上传的文件类型。
但是,当发送.docx文件时,文件类型为:
application/zip
而不是:
application/vnd.openxmlformats-officedocument.wordprocessingml.document
如何更改此行为?
你好,我正在尝试通过finfo_file函数来验证上传的文件类型。
但是,当发送.docx文件时,文件类型为:
application/zip
而不是:
application/vnd.openxmlformats-officedocument.wordprocessingml.document
如何更改此行为?
就我现在而言,供应商特定的文件类型(vnd.)没有标准化(由任何RFC),因此不受file_info()的保护。 是一个压缩的xml格式,这就是原因,为什么返回(完全正确)。您可以解压缩文件并测试结果的哑剧类型,但这将导致文档使用的(也是完全正确的)和其他文件。为了在不同的XML格式之间做出差异,必须分析其内容,并且必须知道它的外观,以及什么。.docx
file_info()
application_zip
xml
file_info()
这在 debian 上是有效的。将此添加到 /etc/magic:
#------------------------------------------------------------------------------
# $File: msooxml,v 1.1 2011/01/25 18:36:19 christos Exp $
# msooxml: file(1) magic for Microsoft Office XML
# From: Ralf Brown <ralf.brown@gmail.com>
# .docx, .pptx, and .xlsx are XML plus other files inside a ZIP
# archive. The first member file is normally "[Content_Types].xml".
# Since MSOOXML doesn't have anything like the uncompressed "mimetype"
# file of ePub or OpenDocument, we'll have to scan for a filename
# which can distinguish between the three types
# start by checking for ZIP local file header signature
0 string PK\003\004
# make sure the first file is correct
>0x1E string [Content_Types].xml
# skip to the second local file header
# since some documents include a 520-byte extra field following the file
# header, we need to scan for the next header
>>(18.l+49) search/2000 PK\003\004
# now skip to the *third* local file header; again, we need to scan due to a
# 520-byte extra field following the file header
>>>&26 search/1000 PK\003\004
# and check the subdirectory name to determine which type of OOXML
# file we have
>>>>&26 string word/ Microsoft Word 2007+
!:mime application/msword
>>>>&26 string ppt/ Microsoft PowerPoint 2007+
!:mime application/vnd.ms-powerpoint
>>>>&26 string xl/ Microsoft Excel 2007+
!:mime application/vnd.ms-excel
>>>>&26 default x Microsoft OOXML
!:strength +10
然后,告诉php使用/etc/magic,因为它是数据库:
$finfo = finfo_open(FILEINFO_MIME,"/etc/magic");