懒人前言:
最终(工作)算法可以在最后的源代码列表中找到。
我之前也记录了这些步骤,因为提问者说“不是程序员”。
(顺便说一句。如果没有前面的步骤,我感觉无法解释最终代码......)
简介
几年前,我在德国计算机杂志 c't 上发现一篇关于 RGB 图像的放大和缩小的文章。这些算法成为我个人图书馆的一部分,我不时使用它们,例如用于在我们的软件中调整图像的大小 - 主要是为了准备好 OpenGL 纹理。
本文的基本思想是考虑源像素(想象为正方形)覆盖目标像素(反之亦然)的空间比率。因此作者区分了放大和缩小。部分覆盖像素的考虑是使用浮点值完成的。
在阅读问题时,我意识到两种特殊情况:
处理位图的要求(由于单色 LCD 输出)
源与目标的宽度和之比为75/16。
比率 75/16 表示 75×75 源像素映射到 16×16 目标像素,例如4.6875×4.6875 源像素到一个目标像素。因此,源图像中存在部分映射到两个甚至四个相邻目标像素的像素。
关于您的特殊要求,我认为在这种特殊情况下应该可以仅使用整数算术来完成。 (根据您的提示(目标平台是嵌入式 CPU),这应该受到欢迎,因为这些通常不提供本机浮点指令。)
掌握 1D
为了热身,我从
字节而不是位
实现了单行图像的缩小:
这个想法是将源像素值累积到 [0,75] 范围内的灰度级,然后使用二进制阈值再次进行二值化。
#include <iostream>
// convenience type for a byte
typedef unsigned char uint8;
// ratio of source image size and destination image size
enum { nR = 75, dR = 16 };
// source image size
enum { wSrc = 1 * nR };
// destination image size
enum { wDst = dR * wSrc / nR };
// binary threshold
enum { tBin = nR / 2 };
// source image
static uint8 imgSrc[wSrc] = {
1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, // 0 ... 15
1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, // 16 ... 31
1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, // 32 ... 47
1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0, 0, 0, 0, 0, 0, // 48 ... 63
1, 1, 1, 1, 1, 1, 1, 1, 0, 0, 0 // 64 ... 74
};
// destination image
static uint8 imgDst[wDst];
// returns a source pixel.
inline int getPixel(int x) { return imgSrc[x]; }
// stores a destination pixel
inline void setPixel(int x, int value)
{
imgDst[x] = !!value; // forces destination value to 0 or 1
}
// prints an image.
void printImg(
int w, // width of image
const uint8 *img) // the image data
{
for (int x = 0; x < w; ++x) std::cout << (char)('0' + img[x]);
std::cout << std::endl;
}
// main function.
int main()
{
// print source image for visual check
std::cout << "Source image (" << wSrc << "):" << std::endl;
printImg(wSrc, imgSrc);
// scale x
int xSrc = 0; int n = 0;
for (int xDst = 0; xDst < wDst; ++xDst) {
int value = 0; // destination pixel accumulator
// process right of cut pixel
if (n) { value += (dR - n) * getPixel(xSrc); ++xSrc; n -= dR; }
n += nR;
// process full pixels
for (; n >= dR; ++xSrc, n -= dR) value += dR * getPixel(xSrc);
// process left of cut pixel
if (n) value += n * getPixel(xSrc);
// store value: 0 ... tBin -> 0, tBin + 1 ... wSrc -> 1
setPixel(xDst, value >= tBin);
}
// print destination image for visual check
std::cout << "Destination image (" << wDst << "):"
<< std::endl;
printImg(wDst, imgDst);
// done
return 0;
}
我在 VisualStudio 2013 中编译测试,得到如下输出:
Source image (75):
111111110000000011111111000000001111111100000000111111110000000011111111000
Destination image (16):
1101100110110010
记住大约 5 个源像素映射到 1 个目标像素,输出看起来对我来说已经足够了。
扩展到二维
下一步是扩展二维图像的第一个样本。我很快意识到我的累积方法必须扩展到完整的目标图像行。这是使用values 数组而不是单个value 来实现的。按照我的第一种方法,必须处理两次拆分的源图像行。为了防止代码重复,我为此引入了辅助函数:accuPixel() 和 accuRow()。
#include <cassert>
#include <iostream>
// convenience type for a byte
typedef unsigned char uint8;
// convenience type for an image
struct Image {
int w, h; // width and height of image
uint8 *data; // image data
int getPixel(int x, int y) const
{
assert(x >= 0 && x < w);
assert(y >= 0 && y < h);
return data[x * w + y];
}
void setPixel(int x, int y, int value)
{
assert(x >= 0 && x < w);
assert(y >= 0 && y < h);
data[x * w + y] = !!value; // '!!' forces dest. value to 0 or 1
}
void print() const
{
for (int y = 0; y < h; ++y) {
for (int x = 0; x < w; ++x) {
std::cout << (char)('0' + data[y * w + x]);
}
std::cout << std::endl;
}
}
};
// ratio of source image size and destination image size
enum { nR = 75, dR = 16 };
// source image size
enum { wSrc = 1 * nR, hSrc = 1 * nR };
// destination image size
enum { wDst = dR * wSrc / nR, hDst = dR * hSrc / nR };
// binary threshold
enum { tBin = nR * nR / 2 };
// source image
static uint8 dataSrc[wSrc * hSrc];
static Image imgSrc = {
/* int w, h: */ wSrc, hSrc,
/* uint8 *data: */ dataSrc
};
// destination image
static uint8 dataDst[wDst * hDst];
static Image imgDst = {
/* int w, h: */ wDst, hDst,
/* uint8 *data: */ dataDst
};
/* accumulates value for a destination pixel from the according number
* of source pixels in one source image row.
*/
void accuPixel(
int &value, // the accumulation value (updated)
const Image &imgSrc, // the source image
int &xSrc, // column index of source pixels (updated)
int ySrc, // row index of source pixels
int &n, // counter of accumulated values (updated)
int fY) // vertical weight of row
{
// process right part of cut pixel
if (n) {
value += fY * (dR - n) * imgSrc.getPixel(xSrc, ySrc);
++xSrc; n -= dR;
}
n += nR;
// process full pixels
for (; n >= dR; ++xSrc, n -= dR) {
value += fY * dR * imgSrc.getPixel(xSrc, ySrc);
}
// process left part of cut pixel
if (n) value += fY * n * imgSrc.getPixel(xSrc, ySrc);
}
/* accumulates values for one destination image row from one source
* image row.
*/
void accuRow(
int wDst, // width of destination image
int *values, // accumulation values for destination row
const Image &imgSrc, // the source image
int ySrc, // row index of source pixels
int fY) // vertical weight of row
{
for (int xSrc = 0, n = 0, xDst = 0; xDst < wDst; ++xDst) {
accuPixel(values[xDst], imgSrc, xSrc, ySrc, n, fY);
}
}
// main function
int main()
{
// fill source image with a chess board pattern
for (int y = 0; y < hSrc; ++y) {
for (int x = 0; x < wSrc; ++x) {
imgSrc.setPixel(x, y, (x % 16 < 8) == (y % 16 < 8));
}
}
// print source image for visual check
std::cout << "Source image (" << wSrc << 'x' << hSrc << "):"
<< std::endl;
imgSrc.print();
// scale source image to destination image
int ySrc = 0; int m = 0;
for (int yDst = 0; yDst < hDst; ++yDst) {
int values[wDst];
for (int &value : values) value = 0; // init accu values
// process bottom of cut row
if (m) {
accuRow(imgDst.w, values, imgSrc, ySrc, dR - m);
++ySrc; m -= dR;
}
m += nR;
// process full rows
for (; m >= dR; ++ySrc, m -= dR) {
accuRow(imgDst.w, values, imgSrc, ySrc, dR);
}
// process top of cut row
if (m) accuRow(imgDst.w, values, imgSrc, ySrc, m);
// process accumulated values
for (int xDst = 0; xDst < wDst; ++xDst) {
imgDst.setPixel(xDst, yDst, values[xDst] >= tBin);
}
}
// print destination image for visual check
std::cout << "Destination image (" << wDst << 'x' << hDst << "):"
<< std::endl;
imgDst.print();
// done
return 0;
}
程序的输出(和输入一样)是一个棋盘。但是,由于插值和随后的二进制分离,输出棋盘格没有大小相等的单元格。
Bit 地图的实际缩小比例
缩放达到我的预期后,示例代码完成:
Image 类已修改为支持位图。如果我使用将值打包为位的std::vector<bool>(std::vector<> 的专用版本),这将很容易。这可能简化了部分代码。我决定反对std::vector<bool>,因为我不确定 OP 中如何提供数据。我相信,我的“显式”C++ 示例代码更容易适应提问者平台上现有的数据模型。
我考虑文件 I/O 以使示例更灵活。我不确定 OP 中的图像格式。我的第一个想法是XMP 只是一个拼写错误,意思是XPM。但后来我开始怀疑并用谷歌搜索了一下。因此,我找到了XMP。可以是这个意思吗?如果我理解正确,XMP 是元数据的标准,可能会添加到某些图像格式,如 JPEG 和 TIFF。所以,我还是不确定...
为了解决这个问题,我决定改用一种文件格式,加载和保存只需要几行代码:PBM。
一旦我实施了 PBM I/O,我就在两个问题上苦苦挣扎,恕我直言,这两个问题值得注意:
如果图像行的长度不是 8 的倍数:行是否字节对齐?因此,我将_bPR // bits per row 成员添加到我的Image 类中。在 PBM 的情况下,行 字节对齐。 (我将带有 GIMP 的 Wikipedia 'J' 示例图像从 ASCII 转换为 RAW 版本以进行检查。)
第一个工作版本(没有崩溃)产生的输出图像看起来并不完全错误,但不知何故“条纹错误”。因此,我得出的结论是,我以错误的顺序存储了每个字节的位。 (从两个可能的解决方案中,我最初选择了错误的一个。)正确的方法是一个字节中最左边的像素必须存储在它的最高有效位中。 (在相反的情况下,Image::getPixel() 和 Image::setPixel() 中的位移必须更改。我将(在 PBM 的情况下)错误版本保留为禁用代码,只是为了这种情况。)
最终的示例代码:
#include <cassert>
#include <iostream>
#include <fstream>
#include <sstream>
#include <string>
// convenience type for bytes
typedef unsigned char uint8;
// image helper class
class Image {
private: // variables:
int _w, _h; // image size
int _bPR; // bits per row
uint8 *_data; // image data
public: // methods:
// constructor.
Image(): _w(0), _h(0), _bPR(0), _data(nullptr) { }
// destructor.
~Image() { free(); }
// returns width of image.
int w() const { return _w; }
// returns height of image.
int h() const { return _h; }
// returns data.
const uint8* data() const { return _data; }
// returns data size (in bytes).
size_t size() const { return (_h * _bPR + 7) / 8; }
// clears image.
void free()
{
delete[] _data; _data = 0; _w = _h = _bPR = 0;
}
// allocates image data.
uint8* alloc( // returns allocated buffer or 0 in case of error
int w, // image width
int h, // image height
int bPR) // bits per row
{
assert(w >= 0 && w <= bPR);
assert(h >= 0);
free();
size_t size = (h * bPR + 7) / 8;
if (size && (_data = new uint8[size])) {
_w = w; _h = h; _bPR = bPR;
}
return _data;
}
// returns pixel.
int getPixel(
int x, // column
int y) // row
const {
assert(x >= 0 && x < _w);
assert(y >= 0 && y < _h);
#if 0 // wrong for PBM
int b = y * _bPR + x, bit = b % 8; // most left pixel is LSB
#else // correct for PBM
int b = y * _bPR + x, bit = 7 - b % 8; // most left pixel is MSB
#endif // 0
return _data[b / 8] >> bit & 1;
}
// sets pixel.
void setPixel(
int x, // column
int y, // row
int value) // value (should be 0 or 1)
{
assert(x >= 0 && x < _w);
assert(y >= 0 && y < _h);
int b = y * _bPR + x;
#if 0 // wrong for PBM
uint8 *pB = _data + b / 8, bit = b % 8; // most left pixel is LSB
#else // correct for PBM
uint8 *pB = _data + b / 8, bit = 7 - b % 8; // most left pixel is MSB
#endif // 0
*pB &= (uint8)~(1 << bit); *pB |= !!value << bit; // bit fiddling
}
};
// reads a PBM binary file.
void readPBM(
std::istream &in, // input stream (to read from)
Image &img) // image to store read data into
{
std::string buffer;
std::getline(in, buffer);
if (buffer != "P4") {
throw "ERROR! File is not a PBM binary file.";
}
do {
std::getline(in, buffer);
} while (buffer[0] == '#');
std::istringstream sIn(buffer);
int w = 0, h = 0;
sIn >> w >> h;
// PBM stores rows aligned to bytes
int bitsPerRow = (w + 7) & ~0x7;
// allocate data memory
char *data = (char*)img.alloc(w, h, bitsPerRow);
// read rest of file at once
in.read(data, img.size());
}
// writes a PBM binary file.
void writePBM(
std::ostream &out, // output stream (to write to)
const Image &img) // image which shall be written
{
out << "P4" << std::endl
<< img.w() << ' ' << img.h() << std::endl;
out.write((const char*)img.data(), img.size());
}
// converts a text to an integer.
int strToI( // returns the integer or throws
const char *text) // text to convert
{
const char *end = text; int value = strtol(text, (char**)&end, 0);
if (end == text || *end != '\0') throw "Not a number.";
return value;
}
/* accumulates value for a destination pixel from the according number
* of source pixels in one source image row.
*/
void accuPixel(
int &value, // the accumulation value (updated)
const Image &imgSrc, // the source image
int &xSrc, // column index of source pixels (updated)
int ySrc, // row index of source pixels
int &n, // counter of accumulated values (updated)
int fY, // vertical weight of row
int nR, // numerator of ratio (source to destination image size)
int dR) // denominator of ratio (source to destination image size)
{
// process right part of cut pixel
if (n) {
value += fY * (dR - n) * imgSrc.getPixel(xSrc, ySrc);
++xSrc; n -= dR;
}
n += nR;
// process full pixels
for (; n >= dR; ++xSrc, n -= dR) {
value += fY * dR * imgSrc.getPixel(xSrc, ySrc);
}
// process left part of cut pixel
if (n) value += fY * n * imgSrc.getPixel(xSrc, ySrc);
}
/* accumulates values for one destination image row from one source
* image row.
*/
void accuRow(
int wDst, // width of destination image
int *values, // accumulation values for destination row
const Image &imgSrc, // the source image
int ySrc, // row index of source pixels
int fY, // vertical weight of row
int nR, // numerator of ratio (source to destination image size)
int dR) // denominator of ratio (source to destination image size)
{
for (int xSrc = 0, n = 0, xDst = 0; xDst < wDst; ++xDst) {
accuPixel(values[xDst], imgSrc, xSrc, ySrc, n, fY, nR, dR);
}
}
// scales source image to destination image.
void scale(
const Image &imgSrc, // source image
Image &imgDst, // destination image
int nR, // numerator of ratio (source to destination image size)
int dR, // denominator of ratio (source to destination image size)
int tBin) // binary threshold e.g. nR * nR / 2
{
// allocate space for destination image
const int wDst = dR * imgSrc.w() / nR;
const int hDst = dR * imgSrc.h() / nR;
if (!imgDst.alloc(wDst, hDst, wDst + 7 & ~7)) {
throw "ERROR! Allocation of destination image failed!";
}
int *values = new int[wDst]; // aux. buffer to accumulate values
for (int ySrc = 0, m = 0, yDst = 0; yDst < hDst; ++yDst) {
// init accu values
for (int i = 0; i < wDst; ++i) values[i] = 0;
// process bottom of cut row
if (m) {
accuRow(wDst, values, imgSrc, ySrc, dR - m, nR, dR);
++ySrc; m -= dR;
}
m += nR;
// process full rows
for (; m >= dR; ++ySrc, m -= dR) {
accuRow(wDst, values, imgSrc, ySrc, dR, nR, dR);
}
// process top of cut row
if (m) accuRow(wDst, values, imgSrc, ySrc, m, nR, dR);
// process accumulated values
for (int xDst = 0; xDst < wDst; ++xDst) {
imgDst.setPixel(xDst, yDst, values[xDst] > tBin);
}
}
delete[] values; // free aux. buffer
}
// main function
int main( // returns 0 on success and another value in error case
int argc, // number of command line arguments
char **argv) // command line arguments
{
// check for sufficient number of arguments
if (argc <= 4) {
std::cerr << "ERROR! Missing command line arguments." << std::endl;
std::cout
<< "Usage:" << std::endl
<< argv[0] << " INFILE OUTFILE NR DR" << std::endl
<< "where" << std::endl
<< "INFILE ... file name of PBM input file (must exist)" << std::endl
<< "OUTFILE ... file name of PBM output file (overwritten if existing)" << std::endl
<< "NR ... numerator of ratio (src. to dest. image size)" << std::endl
<< "DR ... denominator of ratio (src. to dest. image size)" << std::endl
<< "NR and DR must be (not too large) positive integers: 0 < DR < NR" << std::endl;
return 1; // ERROR!
}
try {
// read command line arguments
const char *fileIn = argv[1];
const char *fileOut = argv[2];
int nR;
try {
nR = strToI(argv[3]);
} catch (const char*) {
throw "ERROR in $3! (Not a number.)";
}
int dR;
try {
dR = strToI(argv[4]);
} catch (const char*) {
throw "ERROR in $4! (Not a number.)";
}
int tBin = nR * nR / 2; // might become cmd. line arg. also
// read input file
Image imgSrc;
std::ifstream fIn(fileIn, std::ios::in | std::ios::binary);
fIn.exceptions(std::ifstream::badbit);
readPBM(fIn, imgSrc);
// scale source image to destination image
Image imgDst;
scale(imgSrc, imgDst, nR, dR, tBin);
// write output file
std::ofstream fOut(fileOut, std::ios::out | std::ios::binary);
fOut.exceptions(std::ofstream::badbit);
writePBM(fOut, imgDst);
} catch (const char *error) {
std::cerr << error << std::endl;
return 1; // ERROR!
} catch (const std::exception &error) {
std::cerr << error.what() << std::endl;
return 1; // ERROR!
}
// done (probably successfully)
return 0;
}
为了测试示例代码,我准备了一张我的照片作为示例图像。原图是猫莫里茨在玩螺丝:
我稍微 GIMP 了一下以获得合适的示例图像(主要是因为 GIMP 可以写入、加载和显示 PBM 文件):
虽然我在 VisualStudio 2013 中进行了所有开发和测试,但以下示例会话已使用 g++(在 Windows 10(64 位)上的 cygwin 中)完成:
$ g++ --version
g++ (GCC) 5.4.0
$ g++ -std=c++11 -o scale-bitmap scale-bitmap.cc
$ ./scale-bitmap cat.bin.pbm out.bin.pbm 75 16
$
这产生了以下输出:
如果我没记错的话,示例代码只是实现了Bilinear filtering,这可能是简单地从源图像中删除行和列之后的第二个最糟糕的方法。
如示例输出所示,输出的质量相当有限。更复杂的处理可能会获得更好的结果:
更好的插值可能会有所帮助。维基百科文章Image scaling 和Pixel art scaling algorithms 可能是一个好的开始。
特别是对于单色图像,Dithering 可能是一个选项。
所有这些好东西肯定需要更多的开发工作和代码(如果不在库中使用的话)。
但是,更改二进制阈值tBin 可能会有所改进。我没有尝试过,但我可以想象这是因为我在准备测试图像时使用了 GIMP 中的二进制阈值...
最后但并非最不重要的一点
在写这篇文档的时候。我还发现了一个类似的问题SO: Image downscaling algorithm。 ...并在发送此答案后看到提问者已经提到过...
如果我将nR 和dR 分开用于水平和垂直缩放,该算法也可以应用于非比例缩放。改变它应该不会太难,但在 OP 中不是必需的。
最后,我猜到了目标平台的局限性。关于所描述的 OP 的源图像和目标图像尺寸,最高累积值为 75 * 75 = 5625(将所有源像素缩小为 1 - 一个完整的白色(或黑色?)区域)。这些都是好消息,因为即使 Atmel ATmega 的 C/C++ 编译器仅提供 16 位整数,示例代码也应该可以正常工作。